Chat Completions
POST /v1/chat/completions 接受一段对话,并以 OpenAI Chat Completions 格式返回模型的下一条消息。可从任何 OpenAI SDK 或直接通过 HTTP 使用;本页是逐字段的参考。
POST https://api.shannon-ai.com/v1/chat/completions
最小的请求只包含一个模型 id 和一条用户消息。
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.shannon-ai.com/v1",
)
response = client.chat.completions.create(
model="shannon-3",
messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(response.choices[0].message.content) import OpenAI from "openai";
const client = new OpenAI({
apiKey: "YOUR_API_KEY",
baseURL: "https://api.shannon-ai.com/v1",
});
const response = await client.chat.completions.create({
model: "shannon-3",
messages: [{ role: "user", content: "Say hello in one sentence." }],
});
console.log(response.choices[0].message.content); curl https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "shannon-3",
"messages": [{"role": "user", "content": "Say hello in one sentence."}]
}' 回复是一个 JSON 对象:
{
"id": "chatcmpl-5f0c1e7a9b3d4c62a8e1f07d2b46c9a3",
"object": "chat.completion",
"created": 1791625200,
"model": "shannon-3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello, it is good to meet you.",
"reasoning_content": "The user wants a greeting in one sentence. Keep it short and friendly."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1184,
"completion_tokens": 46,
"total_tokens": 1230
}
} 请求头
请求头
| 请求头 | 值 | 说明 |
|---|---|---|
Authorization | Bearer YOUR_API_KEY | 你的 API 密钥。在每个端点上,也可以用 x-api-key: YOUR_API_KEY 代替。 |
Content-Type | application/json | 必填。其他任何值都会返回 415。 |
x-request-id | 可选。你自己的请求 id。它会原样出现在回复中。 |
回复头
| 请求头 | 说明 |
|---|---|
x-request-id | 每个回复都有,包括错误和流:你发送的值;若未发送,则是 12 位十六进制字符。报告问题时请附上它。 |
content-type | application/json,当 stream 为 true 时为 text/event-stream。 |
请求字段
只有 messages 是必需的。适用模型 一列列出了字段会改变回复的模型。托管开源权重模型即模型列表中的十二个 id;Shannon 3 系列为 shannon-3、shannon-3-pro、shannon-3.1 和 shannon-3.1-pro。 模型与定价
| 字段 | 类型 | 默认值 | 说明 | 适用模型 |
|---|---|---|---|---|
model | string | shannon-1.6-lite | 作答的模型:模型列表中的 id。每个请求都要发送。匹配不区分大小写。未发布的 id 会返回 400 unknown model。 | 所有模型 |
messages | array | 必填。 对话内容,按从旧到新排列。参见下文的消息。 | 所有模型 | |
stream | boolean | false | true 会在回复生成的同时以 server-sent events 形式发送。 | 所有模型 |
max_tokens | integer | 4096 | 回复的上限,以 tokens 计。1 至 65,536 范围之外的值会被调整到该范围内。它同时也是请求运行期间从你的余额中预留的数量。参见下文的输出长度。 | 托管开源权重模型、shannon-1.6-lite、shannon-1.6-pro、shannon-coder-1 |
max_completion_tokens | integer | 与 max_tokens 相同。两者都发送时,使用 max_tokens。 | 托管开源权重模型、shannon-1.6-lite、shannon-1.6-pro、shannon-coder-1 | |
temperature | number | 采样温度。在托管开源权重模型上默认值为 1,取值保持在 0 到 2 之间。 | 托管开源权重模型、shannon-1.6-lite、shannon-1.6-pro、shannon-coder-1 | |
top_p | number | 0.95 | 核采样。取值保持在 0 到 1 之间。 | 托管开源权重模型 |
seed | integer | 采样器的种子,任意整数。若不提供,种子由模型和对话推导而来,因此同一请求发送两次会使用相同的种子。 | 托管开源权重模型 | |
stop | string | array | 字符串或字符串数组。最多使用 4 个。答案会在第一个出现的停止文本之前结束;停止文本本身不会返回。 | 托管开源权重模型 | |
reasoning_effort | string | high | 模型回答前的推理量:off、low、medium 或 high。none 和 minimal 等同于 off,default 等同于 medium,max 等同于 high。其他任何值都会返回 400。 | 托管开源权重模型 |
reasoning | object | 同一设置的对象形式:{"effort": "low"}。两者都发送时,使用 reasoning_effort。 | 托管开源权重模型 | |
tools | array | 模型可以调用的函数,每个形如 {"type": "function", "function": {"name", "description", "parameters"}}。模型的调用会在 tool_calls 中返回;由你的代码运行它们。 | 所有模型 | |
tool_choice | string | object | auto | "auto" 由模型决定。"required" 强制其调用工具。{"type": "function", "function": {"name": "…"}} 强制其调用指定工具。 | 托管开源权重模型 |
response_format | object | {"type": "json_object"} 返回 JSON 答案,{"type": "json_schema", "json_schema": {…}} 返回遵循你的 schema 的答案。 | 所有 Shannon 档位;托管开源权重模型按各 id 的列出情况而定 | |
web_search | boolean | false | true 让模型在回答前先搜索网页。 | shannon-1.6-*、shannon-2-*、Shannon 3 系列 |
其他 OpenAI 字段,如 n、user、stream_options、parallel_tool_calls、presence_penalty、frequency_penalty、logit_bias、logprobs、metadata、store 和 prompt_cache_key,会被接受,以便现有客户端代码无需修改即可运行。它们不会改变回复:始终只有一个 choice,且流始终以用量结尾。
字段的 JSON 类型有误(例如 "max_tokens": "100")会返回 422。缺少 messages 的请求同样如此。
工具、结构化输出、推理和网页搜索各有专页: 函数调用, 结构化输出, 推理强度, 内置网页搜索.
带选项的请求
此请求设置了系统消息、采样字段和推理强度。它使用托管开源权重模型,这些设置都会生效。
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.shannon-ai.com/v1",
)
response = client.chat.completions.create(
model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
messages=[
{"role": "system", "content": "You are a physics teacher. Answer in two sentences."},
{"role": "user", "content": "Why is the sky blue?"},
],
max_tokens=512,
temperature=0.3,
top_p=0.9,
seed=7,
stop=["\n\n"],
reasoning_effort="low",
)
message = response.choices[0].message
print(message.reasoning_content) # the reasoning
print(message.content) # the answer
print(response.usage) import OpenAI from "openai";
const client = new OpenAI({
apiKey: "YOUR_API_KEY",
baseURL: "https://api.shannon-ai.com/v1",
});
const response = await client.chat.completions.create({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
messages: [
{ role: "system", content: "You are a physics teacher. Answer in two sentences." },
{ role: "user", content: "Why is the sky blue?" },
],
max_tokens: 512,
temperature: 0.3,
top_p: 0.9,
seed: 7,
stop: ["\n\n"],
reasoning_effort: "low",
});
const message = response.choices[0].message;
console.log(message.reasoning_content); // the reasoning
console.log(message.content); // the answer
console.log(response.usage); curl https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a physics teacher. Answer in two sentences."},
{"role": "user", "content": "Why is the sky blue?"}
],
"max_tokens": 512,
"temperature": 0.3,
"top_p": 0.9,
"seed": 7,
"stop": ["\n\n"],
"reasoning_effort": "low"
}' 回复的结构与上文相同。在托管开源权重模型上,其 usage 多出两项细节:从缓存读取的 prompt tokens,以及用于推理的 tokens。
{
"usage": {
"prompt_tokens": 31,
"completion_tokens": 62,
"total_tokens": 93,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 21
}
}
} 输出长度
max_tokens 有两个作用。第一,它是请求开始时从你的余额中预留的 tokens 数。回复完成后,该数量会被请求实际用掉的 tokens 取代。如果 max_tokens 大于你余额的剩余量,即使回复本身放得下,请求也会返回 429 Quota exceeded。发送较小的 max_tokens 可以少预留一些。
shannon-coder-1 在此端点上的计数方式不同:每个请求算作你的计划中的一次 Shannon Coder 调用,且不会为其预留 tokens。 限制与余额
第二,它限制这些模型上回复的长度:
| 模型 | max_tokens 对回复长度所起的作用 |
|---|---|
shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 | 回复达到上限时停止。流随后以 finish_reason length 结束。 |
| 托管开源权重模型 | 答案文本在 max_tokens 处停止。推理不计入其中。小于 256 的值按 256 处理。 |
未提供 max_tokens 或 max_completion_tokens 时,值为 4,096。在 shannon-coder-1 上为 65,536。
消息
每条消息是带有 role 和 content 的对象。content 是字符串;当消息不止包含文本时,则是由多个部分组成的数组。
| 角色 | 说明 | 适用模型 |
|---|---|---|
system | 给模型的指令。请放在最前面。在 Shannon 档位上,使用的是第一条 system 消息。 | 托管开源权重模型、shannon-1.6-*、shannon-2-*、shannon-coder-1 |
developer | 按 system 读取。 | 托管开源权重模型 |
user | 你的提问。在 Shannon 档位上,最后一条 user 消息是 prompt,其之前的消息是历史记录。 | 所有模型 |
assistant | 模型先前的回复。在其后发送工具结果时,请保留它的 tool_calls。 | 所有模型 |
tool | 工具调用的结果:tool_call_id 为该调用的 id,content 为字符串形式的结果。 | 所有模型 |
使用 Shannon 3 系列 id 时,请把必须遵守的指令放进 user 消息。
在 Shannon 档位上,既没有用户文本也没有 tools 的请求会返回 400 No user message provided。
内容部分
| 部分 | 说明 | 可用于 |
|---|---|---|
{"type": "text", "text": "…"} | 纯文本。 | 所有模型 |
{"type": "image_url", "image_url": {"url": "…"}} | 图片,以带 base64 内容的 data: URL 或 http(s) URL 形式提供。 | Shannon 3 系列、shannon-1.6-lite、shannon-1.6-pro,以及列出图片输入的托管开源权重模型 |
{"type": "file", "source": {"type": "base64", "media_type": "application/pdf", "data": "…"}} | 一个文档(PDF、Word、PowerPoint 或 Excel 文件),通过 base64 编码或 URL 提供。 | Shannon 3 系列 |
大小、限制和完整的形式列表另有专页说明。 图片与文件
回复对象
| 字段 | 类型 | 说明 |
|---|---|---|
id | string | chatcmpl- 后接 32 位十六进制字符。 |
object | string | 始终为 chat.completion。 |
created | integer | 回复的时间,以 Unix 秒计。 |
model | string | 作答模型的规范 id。其拼写可能与你发送的 id 不同。 |
choices | array | 始终恰好一个 choice,其 index 为 0。 |
choices[0].message.role | string | 始终为 assistant。 |
choices[0].message.content | string | null | 答案文本。带有 tool_calls 时,在 Shannon 档位上为 null;托管开源权重模型可以在调用之外同时发送文本。 |
choices[0].message.reasoning_content | string | null | 模型在答案之前写下的推理内容;没有时为 null。 |
choices[0].message.tool_calls | array | 仅在模型调用工具时出现。每个条目包含 id、type(function)以及 function,其中 name 和 arguments 以 JSON 字符串形式给出。 |
choices[0].message.annotations | array | 仅出现在带 web_search: true 且搜索找到了结果的请求中。content 里的标记所指的每个来源对应一个 url_citation,包含 url、title、start_index 和 end_index(标记的位置,按字符计数,不含结束位置)。 |
choices[0].finish_reason | string | 回复结束的原因。见结束原因。 |
usage | object | 该请求的 tokens。见用量。 |
sources | array | 仅出现在带 web_search: true 且搜索找到了结果的请求中:交给模型的结果,每条包含 index、title 和 url。答案中的 [1] 就是 index 为 1 的条目。 |
结束原因
| finish_reason | 说明 |
|---|---|
stop | 模型完成了回答,或出现了某个 stop 字符串。 |
tool_calls | 模型调用了一个或多个工具。请运行它们,并在 tool 消息中发送结果。 |
length | 回复在输出上限处被截断。会在 shannon-1.6-lite、shannon-1.6-pro、shannon-coder-1 和 Shannon 3 系列的流中报告。 |
非流式回复报告 stop 或 tool_calls。
用量
| 字段 | 类型 | 说明 | 可用于 |
|---|---|---|---|
usage.prompt_tokens | integer | 输入 tokens。 | 所有模型 |
usage.completion_tokens | integer | 输出 tokens:推理、答案和工具调用的总和。 | 所有模型 |
usage.total_tokens | integer | prompt_tokens 加 completion_tokens。 | 所有模型 |
usage.prompt_tokens_details.cached_tokens | integer | prompt_tokens 中从 prompt 缓存读取的部分。 | 托管开源权重模型 |
usage.completion_tokens_details.reasoning_tokens | integer | completion_tokens 中用于推理的部分。 | 托管开源权重模型 |
在托管开源权重模型上,prompt_tokens 是用该模型自己的分词器统计的消息和工具定义,加上任何图片的 tokens。tokens 计数端点会在你发送之前返回相同的数字。 Token 计数
在 Shannon 档位上,prompt_tokens 统计模型为写出回复所读取的全部内容,因此比你的消息文本本身要大。
流式传输
将 stream 设为 true 时,回复以 chat.completion.chunk 事件到达,并以 data: [DONE] 结束。在它之前的最后一个数据块带有 finish_reason 和 usage;无需 stream_options。数据块的结构、保活行以及流内的错误另有专页说明。 流式
错误
错误是带 error 成员的 JSON 对象。检查按以下顺序进行:API 密钥、请求体、模型 id,然后是余额。下表列出此端点最常返回的错误。完整列表及重试建议见专页。 错误处理
{
"error": {
"type": "invalid_request_error",
"message": "unknown model: no-such-model"
}
} | 状态码 | 类型 | 消息 | 时机 |
|---|---|---|---|
401 | authentication_error | Missing authenticationInvalid API key | 未发送 API 密钥,或密钥未知或已被撤销。 |
400 | invalid_request_error | unknown model: <id> | model 不是已发布的 id。 |
400 | invalid_request_error | No user message provided | Shannon 档位:请求既没有用户文本,也没有 tools。 |
400 | invalid_request_error | <id> does not accept image input | 向不支持图片输入的托管开源权重模型发送了图片部分。 |
400 | invalid_request_error | <id> does not accept response_format | 向不支持结构化输出的托管开源权重模型发送了 response_format。 |
400 | invalid_request_error | unknown reasoning effort '<value>'; expected off, low, medium or high | reasoning_effort 的值不在列表内。 |
422 | invalid_request_error | Failed to deserialize the JSON body into the target type: … | 缺少 messages,或某个字段的 JSON 类型有误。 |
429 | rate_limit_error | Quota exceeded. Upgrade your plan at shannon-ai.com/plan | max_tokens 大于你余额的剩余量。 |
429 | rate_limit_error | Too many requests. Retry in <n>s. | 限流保护:你的账户在一分钟内发送了超过 120 个请求。 |
500 | server_error | The model backend failed to answer. Please retry. | 模型未生成回复。请重新发送请求。 |
502 | api_error | The model backend failed to answer. Please retry. | 同上,适用于 Shannon 3 系列和托管开源权重模型。 |