Skip to content
Overview

Overview

The map of the API: every endpoint, what a request and an error look like, how calls are paid for, and what to know when you come from an OpenAI or Anthropic SDK.

Endpoints

Every endpoint lives under one base URL and is served over HTTPS.

Base URL
https://api.shannon-ai.com
Endpoint Format What it is for
POST /v1/chat/completions OpenAI Chat Completions Send a conversation, get the next answer. With or without streaming.
POST /v1/messages Anthropic Messages The same, in the request and reply shapes of Anthropic SDKs.
POST /v1/responses OpenAI Responses The same, in the Responses shapes. The endpoint keeps no state: send the conversation with every request.
GET /v1/models OpenAI model list List the models with context window, prices and capabilities. Needs no key.
POST /v1/tokenize Shannon API Count the tokens of a text or of a chat request for a hosted open-weight model. Free.
POST /v1/messages/count_tokens Anthropic token count Count the input tokens of a Messages request for a hosted open-weight model. Free.

The three endpoints that produce text reach the same models. Choose the one whose format your code already uses.

Request basics

Header Description
Authorization: Bearer <key> Your API key. Required on every endpoint except GET /v1/models, unless you send x-api-key.
x-api-key: <key> The same key in the header Anthropic SDKs send. Read on every endpoint.
Content-Type: application/json Required on every POST. Without it the reply is 415.
x-request-id: <your id> Optional. Your own id for the request; it comes back in the reply header x-request-id. Without it the API creates one of 12 hexadecimal characters.
  • The body of every POST is one JSON object, up to 32 MiB.
  • A field the API does not know causes no error and has no effect. A request written for another provider does not fail because of an extra field.
  • A known field with the wrong JSON type, or a missing required field, is answered with 422. A body that is not valid JSON is answered with 400.
  • model is one of the ids on Models & pricing. Upper and lower case do not matter.

A reply is JSON, or a stream of server-sent events when the request sets stream to true. Each endpoint answers in its own format. Every reply has the header x-request-id.

What a request passes

A request is checked in a fixed order before a model runs. The first check that fails answers, so a 401 tells you nothing yet about the body.

Error shape

An error is a JSON object with an error that holds type and message. /v1/messages wraps it the way Anthropic SDKs expect; every other path uses the OpenAI shape.

{
  "error": {
    "type": "invalid_request_error",
    "message": "unknown model: gpt-4o"
  }
}
  • Read type and message. code and param are present on some errors only: treat them as optional. param is always null.
  • After a stream has started, the status is already 200. A failure then arrives as an error frame inside the stream.
  • Every error reply carries the header x-request-id.
Status Type When
400 invalid_request_error The body is not valid JSON, the model id is unknown, or the model does not take a kind of input you sent.
401 authentication_error The key is missing or not valid.
404 not_found_error The path does not exist.
405 api_error The path exists, the method is wrong.
413 invalid_request_error The body is larger than 32 MiB.
415 invalid_request_error Content-Type is not application/json.
422 invalid_request_error A field has the wrong JSON type or a required field is missing.
429 rate_limit_error The balance does not cover the request, more than 120 requests arrived in a minute, the Shannon Coder calls of the window are used up, or the model is busy. The message says which.
5xx api_error Status 500, 502, 503 or 504: the request was valid and could not be answered. Send it again. A 500 can carry the type server_error.

Error Handling

Billing and balance

  • There is one balance per account, and chat and API share it: today's plan allowance first, then purchased credit. The API has no quota of its own.
  • A request reserves its output budget (max_tokens, default 4,096) and is then charged for the tokens it really used, at the price of the model.
  • Every reply reports its token counts in usage. The Keys & usage page shows the balance and what each request cost.
  • Every request is served equally. The only limit on request rate is flood protection: 120 requests per minute per account. Requests sent in parallel wait in line.

Limits and balance Models & pricing Keys & usage

Fields that depend on the model

Every model takes the same request. A few fields take effect on some models only; the table names where. The endpoint pages list every field.

Field Description Applied by
system Instructions for the model: a system message on Chat Completions, system on Messages, instructions on Responses. Hosted open-weight models, shannon-1.6-*, shannon-2-*, shannon-coder-1
temperature Sampling temperature. Hosted open-weight models, shannon-1.6-*, shannon-coder-1
top_p Nucleus sampling. Hosted open-weight models
seed A fixed seed for sampling. Hosted open-weight models
stop Up to 4 stop sequences. Hosted open-weight models
reasoning_effort How much the model reasons before it answers. reasoning.effort on Responses, thinking on Messages. Hosted open-weight models
web_search true lets the model search the web for this request. A field of this API, on Chat Completions and Messages. Shannon models except shannon-coder-1
max_tokens The output budget. On every model it sets the amount reserved from your balance. As the limit on the length of the answer: hosted open-weight models, shannon-1.6-*, shannon-coder-1

Chat Completions

Coming from an OpenAI SDK

  • Set the base URL to https://api.shannon-ai.com/v1 and the key to your Shannon key. Chat Completions and Responses calls then work with the SDK as it is.
  • model must be a Shannon id. A model name of another provider, such as gpt-4o, is answered with 400 and unknown model.
  • Reasoning comes in a field of its own: reasoning_content beside content, in the message and in the stream deltas.
  • A stream always carries usage in its last chunk, together with finish_reason.
  • A tool call in a stream arrives as one chunk with the complete arguments string.
  • A reply has one choice.
  • Paths of the OpenAI API that are not in the table above, such as /v1/embeddings, are answered with 404.

Coming from an Anthropic SDK

  • Set the base URL to https://api.shannon-ai.com, without /v1, and the key to your Shannon key. The SDK sends it as x-api-key.
  • model must be a Shannon id.
  • max_tokens is optional on this API. Its default is 4,096.
  • A reply holds content blocks of type thinking, text and tool_use. The first block is not always the text: pick blocks by type.
  • stop_reason is end_turn or tool_use. A stream of a Shannon model can also end with max_tokens.
  • anthropic-version and anthropic-beta are accepted, so the SDK works unchanged. A request does not need them.
  • Errors on /v1/messages have the Anthropic shape: {"type": "error", "error": {…}}.

Coding tools that speak these formats are set up the same way: base URL, key, and a Shannon id as the model. CLI coding tools