Chat Completions
POST /v1/chat/completions takes a conversation and returns the model's next message in the OpenAI Chat Completions format. Use it from any OpenAI SDK or over plain HTTP; this page is the field-by-field reference.
POST https://api.shannon-ai.com/v1/chat/completions
The smallest request is a model id and one user message.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.shannon-ai.com/v1",
)
response = client.chat.completions.create(
model="shannon-3",
messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(response.choices[0].message.content) import OpenAI from "openai";
const client = new OpenAI({
apiKey: "YOUR_API_KEY",
baseURL: "https://api.shannon-ai.com/v1",
});
const response = await client.chat.completions.create({
model: "shannon-3",
messages: [{ role: "user", content: "Say hello in one sentence." }],
});
console.log(response.choices[0].message.content); curl https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "shannon-3",
"messages": [{"role": "user", "content": "Say hello in one sentence."}]
}' The reply is one JSON object:
{
"id": "chatcmpl-5f0c1e7a9b3d4c62a8e1f07d2b46c9a3",
"object": "chat.completion",
"created": 1791625200,
"model": "shannon-3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello, it is good to meet you.",
"reasoning_content": "The user wants a greeting in one sentence. Keep it short and friendly."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1184,
"completion_tokens": 46,
"total_tokens": 1230
}
} Headers
Request headers
| Header | Value | Description |
|---|---|---|
Authorization | Bearer YOUR_API_KEY | Your API key. x-api-key: YOUR_API_KEY is accepted in its place on every endpoint. |
Content-Type | application/json | Required. Any other value returns 415. |
x-request-id | Optional. Your own id for the request. It comes back unchanged on the reply. |
Reply headers
| Header | Description |
|---|---|
x-request-id | On every reply, errors and streams included: the value you sent, or 12 hexadecimal characters when you sent none. Quote it when you report a problem. |
content-type | application/json, or text/event-stream when stream is true. |
Request fields
Only messages is required. The Applied by column names the models on which a field changes the reply. The hosted open-weight models are the twelve ids of the model list; the Shannon 3 family is shannon-3, shannon-3-pro, shannon-3.1 and shannon-3.1-pro. Models & pricing
| Field | Type | Default | Description | Applied by |
|---|---|---|---|---|
model | string | shannon-1.6-lite | The model that answers: an id from the model list. Send it with every request. Matching is not case-sensitive. An id that is not published returns 400 unknown model. | All models |
messages | array | Required. The conversation, oldest message first. See Messages below. | All models | |
stream | boolean | false | true sends the reply as server-sent events while it is written. | All models |
max_tokens | integer | 4096 | Upper limit of the reply, in tokens. A value outside 1 to 65,536 is moved into that range. It is also the amount set aside from your balance while the request runs. See Output length below. | Hosted open-weight models, shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 |
max_completion_tokens | integer | Same as max_tokens. When both are sent, max_tokens is used. | Hosted open-weight models, shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 | |
temperature | number | Sampling temperature. On the hosted open-weight models the default is 1 and values are kept between 0 and 2. | Hosted open-weight models, shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 | |
top_p | number | 0.95 | Nucleus sampling. Values are kept between 0 and 1. | Hosted open-weight models |
seed | integer | Seed of the sampler, any integer. Without it, the seed is derived from the model and the conversation, so the same request sent twice uses the same seed. | Hosted open-weight models | |
stop | string | array | A string or an array of strings. Up to 4 are used. The answer ends before the first one that appears; the stop text itself is not returned. | Hosted open-weight models | |
reasoning_effort | string | high | How much the model reasons before it answers: off, low, medium or high. none and minimal mean off, default means medium, max means high. Any other value returns 400. | Hosted open-weight models |
reasoning | object | The same setting in object form: {"effort": "low"}. When both are sent, reasoning_effort is used. | Hosted open-weight models | |
tools | array | The functions the model may call, each as {"type": "function", "function": {"name", "description", "parameters"}}. The model's calls come back in tool_calls; your code runs them. | All models | |
tool_choice | string | object | auto | "auto" lets the model decide. "required" makes it call a tool. {"type": "function", "function": {"name": "…"}} makes it call that tool. | Hosted open-weight models |
response_format | object | {"type": "json_object"} for a JSON answer, or {"type": "json_schema", "json_schema": {…}} for an answer that follows your schema. | All Shannon tiers; hosted open-weight models as listed per id | |
web_search | boolean | false | true lets the model search the web before it answers. | shannon-1.6-*, shannon-2-*, Shannon 3 family |
Other OpenAI fields, such as n, user, stream_options, parallel_tool_calls, presence_penalty, frequency_penalty, logit_bias, logprobs, metadata, store and prompt_cache_key, are accepted so that existing client code runs unchanged. They do not change the reply: there is always one choice, and a stream always ends with usage.
A field with the wrong JSON type, for example "max_tokens": "100", returns 422. A request without messages does too.
Tools, structured output, reasoning and web search each have their own page: Function Calling, Structured Outputs, Reasoning effort, Built-in Web Search.
A request with options
This request sets a system message, the sampling fields and the reasoning effort. It uses a hosted open-weight model, which applies all of them.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.shannon-ai.com/v1",
)
response = client.chat.completions.create(
model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
messages=[
{"role": "system", "content": "You are a physics teacher. Answer in two sentences."},
{"role": "user", "content": "Why is the sky blue?"},
],
max_tokens=512,
temperature=0.3,
top_p=0.9,
seed=7,
stop=["\n\n"],
reasoning_effort="low",
)
message = response.choices[0].message
print(message.reasoning_content) # the reasoning
print(message.content) # the answer
print(response.usage) import OpenAI from "openai";
const client = new OpenAI({
apiKey: "YOUR_API_KEY",
baseURL: "https://api.shannon-ai.com/v1",
});
const response = await client.chat.completions.create({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
messages: [
{ role: "system", content: "You are a physics teacher. Answer in two sentences." },
{ role: "user", content: "Why is the sky blue?" },
],
max_tokens: 512,
temperature: 0.3,
top_p: 0.9,
seed: 7,
stop: ["\n\n"],
reasoning_effort: "low",
});
const message = response.choices[0].message;
console.log(message.reasoning_content); // the reasoning
console.log(message.content); // the answer
console.log(response.usage); curl https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a physics teacher. Answer in two sentences."},
{"role": "user", "content": "Why is the sky blue?"}
],
"max_tokens": 512,
"temperature": 0.3,
"top_p": 0.9,
"seed": 7,
"stop": ["\n\n"],
"reasoning_effort": "low"
}' The reply has the same shape as above. Its usage adds two details on the hosted open-weight models: the prompt tokens read from cache and the tokens spent on reasoning.
{
"usage": {
"prompt_tokens": 31,
"completion_tokens": 62,
"total_tokens": 93,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 21
}
}
} Output length
max_tokens does two things. First, it is the number of tokens set aside from your balance when the request starts. When the reply is complete, that amount is replaced by the tokens the request used. If max_tokens is larger than what is left of your balance, the request returns 429 Quota exceeded even when the reply itself would have fitted. Send a lower max_tokens to set aside less.
shannon-coder-1 is counted differently on this endpoint: each request is one of your plan's Shannon Coder calls, and no tokens are set aside for it. Limits and balance
Second, it limits the length of the reply on these models:
| Models | What max_tokens does |
|---|---|
shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 | The reply stops when it reaches the limit. A stream then ends with finish_reason length. |
| Hosted open-weight models | The answer text stops at max_tokens. Reasoning is not counted against it. Values below 256 act as 256. |
Without max_tokens or max_completion_tokens, the value is 4,096. On shannon-coder-1 it is 65,536.
Messages
Each message is an object with a role and a content. content is a string, or an array of parts when the message carries more than text.
| Role | Description | Applied by |
|---|---|---|
system | Instructions for the model. Put it first. On the Shannon tiers the first system message is the one that is used. | Hosted open-weight models, shannon-1.6-*, shannon-2-*, shannon-coder-1 |
developer | Read as system. | Hosted open-weight models |
user | What you ask. On the Shannon tiers the last user message is the prompt and the messages before it are the history. | All models |
assistant | Earlier replies of the model. Keep its tool_calls when you send a tool result after it. | All models |
tool | The result of a tool call: tool_call_id holds the id of the call and content the result as a string. | All models |
With a Shannon 3 family id, put instructions that must hold into the user message.
On the Shannon tiers a request with no user text and no tools returns 400 No user message provided.
Content parts
| Part | Description | Available on |
|---|---|---|
{"type": "text", "text": "…"} | Plain text. | All models |
{"type": "image_url", "image_url": {"url": "…"}} | An image, as a data: URL with base64 content or as an http(s) URL. | Shannon 3 family, shannon-1.6-lite, shannon-1.6-pro, and the hosted open-weight models that list image input |
{"type": "file", "source": {"type": "base64", "media_type": "application/pdf", "data": "…"}} | A document (PDF, Word, PowerPoint or Excel), as base64 or by URL. | Shannon 3 family |
Sizes, limits and the full list of forms have their own page. Images and files
The reply object
| Field | Type | Description |
|---|---|---|
id | string | chatcmpl- followed by 32 hexadecimal characters. |
object | string | Always chat.completion. |
created | integer | Time of the reply, in Unix seconds. |
model | string | The canonical id of the model that answered. It can differ in spelling from the id you sent. |
choices | array | Always exactly one choice, with index 0. |
choices[0].message.role | string | Always assistant. |
choices[0].message.content | string | null | The answer text. With tool_calls it is null on the Shannon tiers; the hosted open-weight models can send text beside the calls. |
choices[0].message.reasoning_content | string | null | The reasoning the model wrote before the answer, or null when there is none. |
choices[0].message.tool_calls | array | Present only when the model calls tools. Each entry has an id, type function, and function with the name and the arguments as a JSON string. |
choices[0].message.annotations | array | Only on a request with web_search: true whose search found something. One url_citation for each source a marker in content names, with url, title, start_index and end_index (the position of the marker, counted in characters, end not included). |
choices[0].finish_reason | string | Why the reply ended. See Finish reasons. |
usage | object | The tokens of the request. See Usage. |
sources | array | Only on a request with web_search: true whose search found something: the results the model was given, each with index, title and url. [1] in the answer is the entry with index 1. |
Finish reasons
| finish_reason | Description |
|---|---|
stop | The model finished its answer, or a stop string appeared. |
tool_calls | The model calls one or more tools. Run them and send the results in tool messages. |
length | The reply was cut at the output limit. Reported in streams of shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 and the Shannon 3 family. |
A reply that is not streamed reports stop or tool_calls.
Usage
| Field | Type | Description | Available on |
|---|---|---|---|
usage.prompt_tokens | integer | Input tokens. | All models |
usage.completion_tokens | integer | Output tokens: reasoning, answer and tool calls together. | All models |
usage.total_tokens | integer | prompt_tokens plus completion_tokens. | All models |
usage.prompt_tokens_details.cached_tokens | integer | The part of prompt_tokens that was read from the prompt cache. | Hosted open-weight models |
usage.completion_tokens_details.reasoning_tokens | integer | The part of completion_tokens that was spent on reasoning. | Hosted open-weight models |
On the hosted open-weight models, prompt_tokens is your messages and tool definitions counted with the model's own tokenizer, plus the tokens of any images. The token counting endpoints return the same number before you send. Token counting
On the Shannon tiers, prompt_tokens counts everything the model read to write the reply, so it is larger than the text of your messages alone.
Streaming
With stream set to true the reply arrives as chat.completion.chunk events and ends with data: [DONE]. The last chunk before it carries finish_reason and usage; no stream_options are needed. The chunk shapes, keep-alive lines and errors inside a stream have their own page. Streaming
Errors
An error is a JSON object with an error member. Checks run in this order: API key, request body, model id, then balance. The table lists what this endpoint returns most often. The full list, with what to retry, has its own page. Error Handling
{
"error": {
"type": "invalid_request_error",
"message": "unknown model: no-such-model"
}
} | Status | Type | Message | When |
|---|---|---|---|
401 | authentication_error | Missing authenticationInvalid API key | No API key was sent, or the key is unknown or revoked. |
400 | invalid_request_error | unknown model: <id> | model is not a published id. |
400 | invalid_request_error | No user message provided | Shannon tiers: the request has no user text and no tools. |
400 | invalid_request_error | <id> does not accept image input | An image part was sent to a hosted open-weight model without image input. |
400 | invalid_request_error | <id> does not accept response_format | response_format was sent to a hosted open-weight model without structured output. |
400 | invalid_request_error | unknown reasoning effort '<value>'; expected off, low, medium or high | reasoning_effort holds a value outside the list. |
422 | invalid_request_error | Failed to deserialize the JSON body into the target type: … | messages is missing, or a field has the wrong JSON type. |
429 | rate_limit_error | Quota exceeded. Upgrade your plan at shannon-ai.com/plan | max_tokens is larger than what is left of your balance. |
429 | rate_limit_error | Too many requests. Retry in <n>s. | Flood protection: more than 120 requests in one minute on your account. |
500 | server_error | The model backend failed to answer. Please retry. | The model did not produce a reply. Send the request again. |
502 | api_error | The model backend failed to answer. Please retry. | The same, on the Shannon 3 family and the hosted open-weight models. |