Skip to content
Chat Completions

Chat Completions

POST /v1/chat/completions takes a conversation and returns the model's next message in the OpenAI Chat Completions format. Use it from any OpenAI SDK or over plain HTTP; this page is the field-by-field reference.

POST https://api.shannon-ai.com/v1/chat/completions

The smallest request is a model id and one user message.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.shannon-ai.com/v1",
)

response = client.chat.completions.create(
    model="shannon-3",
    messages=[{"role": "user", "content": "Say hello in one sentence."}],
)

print(response.choices[0].message.content)

The reply is one JSON object:

200 JSON
{
  "id": "chatcmpl-5f0c1e7a9b3d4c62a8e1f07d2b46c9a3",
  "object": "chat.completion",
  "created": 1791625200,
  "model": "shannon-3",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello, it is good to meet you.",
        "reasoning_content": "The user wants a greeting in one sentence. Keep it short and friendly."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 1184,
    "completion_tokens": 46,
    "total_tokens": 1230
  }
}

Headers

Request headers

Header Value Description
Authorization Bearer YOUR_API_KEY Your API key. x-api-key: YOUR_API_KEY is accepted in its place on every endpoint.
Content-Type application/json Required. Any other value returns 415.
x-request-id Optional. Your own id for the request. It comes back unchanged on the reply.

Reply headers

Header Description
x-request-id On every reply, errors and streams included: the value you sent, or 12 hexadecimal characters when you sent none. Quote it when you report a problem.
content-type application/json, or text/event-stream when stream is true.

Request fields

Only messages is required. The Applied by column names the models on which a field changes the reply. The hosted open-weight models are the twelve ids of the model list; the Shannon 3 family is shannon-3, shannon-3-pro, shannon-3.1 and shannon-3.1-pro. Models & pricing

Field Type Default Description Applied by
model string shannon-1.6-lite The model that answers: an id from the model list. Send it with every request. Matching is not case-sensitive. An id that is not published returns 400 unknown model. All models
messages array Required. The conversation, oldest message first. See Messages below. All models
stream boolean false true sends the reply as server-sent events while it is written. All models
max_tokens integer 4096 Upper limit of the reply, in tokens. A value outside 1 to 65,536 is moved into that range. It is also the amount set aside from your balance while the request runs. See Output length below. Hosted open-weight models, shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1
max_completion_tokens integer Same as max_tokens. When both are sent, max_tokens is used. Hosted open-weight models, shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1
temperature number Sampling temperature. On the hosted open-weight models the default is 1 and values are kept between 0 and 2. Hosted open-weight models, shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1
top_p number 0.95 Nucleus sampling. Values are kept between 0 and 1. Hosted open-weight models
seed integer Seed of the sampler, any integer. Without it, the seed is derived from the model and the conversation, so the same request sent twice uses the same seed. Hosted open-weight models
stop string | array A string or an array of strings. Up to 4 are used. The answer ends before the first one that appears; the stop text itself is not returned. Hosted open-weight models
reasoning_effort string high How much the model reasons before it answers: off, low, medium or high. none and minimal mean off, default means medium, max means high. Any other value returns 400. Hosted open-weight models
reasoning object The same setting in object form: {"effort": "low"}. When both are sent, reasoning_effort is used. Hosted open-weight models
tools array The functions the model may call, each as {"type": "function", "function": {"name", "description", "parameters"}}. The model's calls come back in tool_calls; your code runs them. All models
tool_choice string | object auto "auto" lets the model decide. "required" makes it call a tool. {"type": "function", "function": {"name": "…"}} makes it call that tool. Hosted open-weight models
response_format object {"type": "json_object"} for a JSON answer, or {"type": "json_schema", "json_schema": {…}} for an answer that follows your schema. All Shannon tiers; hosted open-weight models as listed per id
web_search boolean false true lets the model search the web before it answers. shannon-1.6-*, shannon-2-*, Shannon 3 family

Other OpenAI fields, such as n, user, stream_options, parallel_tool_calls, presence_penalty, frequency_penalty, logit_bias, logprobs, metadata, store and prompt_cache_key, are accepted so that existing client code runs unchanged. They do not change the reply: there is always one choice, and a stream always ends with usage.

A field with the wrong JSON type, for example "max_tokens": "100", returns 422. A request without messages does too.

Tools, structured output, reasoning and web search each have their own page: Function Calling, Structured Outputs, Reasoning effort, Built-in Web Search.

A request with options

This request sets a system message, the sampling fields and the reasoning effort. It uses a hosted open-weight model, which applies all of them.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.shannon-ai.com/v1",
)

response = client.chat.completions.create(
    model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
    messages=[
        {"role": "system", "content": "You are a physics teacher. Answer in two sentences."},
        {"role": "user", "content": "Why is the sky blue?"},
    ],
    max_tokens=512,
    temperature=0.3,
    top_p=0.9,
    seed=7,
    stop=["\n\n"],
    reasoning_effort="low",
)

message = response.choices[0].message
print(message.reasoning_content)  # the reasoning
print(message.content)            # the answer
print(response.usage)

The reply has the same shape as above. Its usage adds two details on the hosted open-weight models: the prompt tokens read from cache and the tokens spent on reasoning.

200 JSON
{
  "usage": {
    "prompt_tokens": 31,
    "completion_tokens": 62,
    "total_tokens": 93,
    "prompt_tokens_details": {
      "cached_tokens": 0
    },
    "completion_tokens_details": {
      "reasoning_tokens": 21
    }
  }
}

Output length

max_tokens does two things. First, it is the number of tokens set aside from your balance when the request starts. When the reply is complete, that amount is replaced by the tokens the request used. If max_tokens is larger than what is left of your balance, the request returns 429 Quota exceeded even when the reply itself would have fitted. Send a lower max_tokens to set aside less.

shannon-coder-1 is counted differently on this endpoint: each request is one of your plan's Shannon Coder calls, and no tokens are set aside for it. Limits and balance

Second, it limits the length of the reply on these models:

Models What max_tokens does
shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 The reply stops when it reaches the limit. A stream then ends with finish_reason length.
Hosted open-weight models The answer text stops at max_tokens. Reasoning is not counted against it. Values below 256 act as 256.

Without max_tokens or max_completion_tokens, the value is 4,096. On shannon-coder-1 it is 65,536.

Messages

Each message is an object with a role and a content. content is a string, or an array of parts when the message carries more than text.

Role Description Applied by
system Instructions for the model. Put it first. On the Shannon tiers the first system message is the one that is used. Hosted open-weight models, shannon-1.6-*, shannon-2-*, shannon-coder-1
developer Read as system. Hosted open-weight models
user What you ask. On the Shannon tiers the last user message is the prompt and the messages before it are the history. All models
assistant Earlier replies of the model. Keep its tool_calls when you send a tool result after it. All models
tool The result of a tool call: tool_call_id holds the id of the call and content the result as a string. All models

With a Shannon 3 family id, put instructions that must hold into the user message.

On the Shannon tiers a request with no user text and no tools returns 400 No user message provided.

Content parts

Part Description Available on
{"type": "text", "text": "…"} Plain text. All models
{"type": "image_url", "image_url": {"url": "…"}} An image, as a data: URL with base64 content or as an http(s) URL. Shannon 3 family, shannon-1.6-lite, shannon-1.6-pro, and the hosted open-weight models that list image input
{"type": "file", "source": {"type": "base64", "media_type": "application/pdf", "data": "…"}} A document (PDF, Word, PowerPoint or Excel), as base64 or by URL. Shannon 3 family

Sizes, limits and the full list of forms have their own page. Images and files

The reply object

Field Type Description
id string chatcmpl- followed by 32 hexadecimal characters.
object string Always chat.completion.
created integer Time of the reply, in Unix seconds.
model string The canonical id of the model that answered. It can differ in spelling from the id you sent.
choices array Always exactly one choice, with index 0.
choices[0].message.role string Always assistant.
choices[0].message.content string | null The answer text. With tool_calls it is null on the Shannon tiers; the hosted open-weight models can send text beside the calls.
choices[0].message.reasoning_content string | null The reasoning the model wrote before the answer, or null when there is none.
choices[0].message.tool_calls array Present only when the model calls tools. Each entry has an id, type function, and function with the name and the arguments as a JSON string.
choices[0].message.annotations array Only on a request with web_search: true whose search found something. One url_citation for each source a marker in content names, with url, title, start_index and end_index (the position of the marker, counted in characters, end not included).
choices[0].finish_reason string Why the reply ended. See Finish reasons.
usage object The tokens of the request. See Usage.
sources array Only on a request with web_search: true whose search found something: the results the model was given, each with index, title and url. [1] in the answer is the entry with index 1.

Finish reasons

finish_reason Description
stop The model finished its answer, or a stop string appeared.
tool_calls The model calls one or more tools. Run them and send the results in tool messages.
length The reply was cut at the output limit. Reported in streams of shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 and the Shannon 3 family.

A reply that is not streamed reports stop or tool_calls.

Usage

Field Type Description Available on
usage.prompt_tokens integer Input tokens. All models
usage.completion_tokens integer Output tokens: reasoning, answer and tool calls together. All models
usage.total_tokens integer prompt_tokens plus completion_tokens. All models
usage.prompt_tokens_details.cached_tokens integer The part of prompt_tokens that was read from the prompt cache. Hosted open-weight models
usage.completion_tokens_details.reasoning_tokens integer The part of completion_tokens that was spent on reasoning. Hosted open-weight models

On the hosted open-weight models, prompt_tokens is your messages and tool definitions counted with the model's own tokenizer, plus the tokens of any images. The token counting endpoints return the same number before you send. Token counting

On the Shannon tiers, prompt_tokens counts everything the model read to write the reply, so it is larger than the text of your messages alone.

Streaming

With stream set to true the reply arrives as chat.completion.chunk events and ends with data: [DONE]. The last chunk before it carries finish_reason and usage; no stream_options are needed. The chunk shapes, keep-alive lines and errors inside a stream have their own page. Streaming

Errors

An error is a JSON object with an error member. Checks run in this order: API key, request body, model id, then balance. The table lists what this endpoint returns most often. The full list, with what to retry, has its own page. Error Handling

400 JSON
{
  "error": {
    "type": "invalid_request_error",
    "message": "unknown model: no-such-model"
  }
}
Status Type Message When
401 authentication_error Missing authentication
Invalid API key
No API key was sent, or the key is unknown or revoked.
400 invalid_request_error unknown model: <id> model is not a published id.
400 invalid_request_error No user message provided Shannon tiers: the request has no user text and no tools.
400 invalid_request_error <id> does not accept image input An image part was sent to a hosted open-weight model without image input.
400 invalid_request_error <id> does not accept response_format response_format was sent to a hosted open-weight model without structured output.
400 invalid_request_error unknown reasoning effort '<value>'; expected off, low, medium or high reasoning_effort holds a value outside the list.
422 invalid_request_error Failed to deserialize the JSON body into the target type: … messages is missing, or a field has the wrong JSON type.
429 rate_limit_error Quota exceeded. Upgrade your plan at shannon-ai.com/plan max_tokens is larger than what is left of your balance.
429 rate_limit_error Too many requests. Retry in <n>s. Flood protection: more than 120 requests in one minute on your account.
500 server_error The model backend failed to answer. Please retry. The model did not produce a reply. Send the request again.
502 api_error The model backend failed to answer. Please retry. The same, on the Shannon 3 family and the hosted open-weight models.