跳到正文
Chat Completions

Chat Completions

POST /v1/chat/completions 接受一段对话,并以 OpenAI Chat Completions 格式返回模型的下一条消息。可从任何 OpenAI SDK 或直接通过 HTTP 使用;本页是逐字段的参考。

POST https://api.shannon-ai.com/v1/chat/completions

最小的请求只包含一个模型 id 和一条用户消息。

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.shannon-ai.com/v1",
)

response = client.chat.completions.create(
    model="shannon-3",
    messages=[{"role": "user", "content": "Say hello in one sentence."}],
)

print(response.choices[0].message.content)

回复是一个 JSON 对象:

200 JSON
{
  "id": "chatcmpl-5f0c1e7a9b3d4c62a8e1f07d2b46c9a3",
  "object": "chat.completion",
  "created": 1791625200,
  "model": "shannon-3",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello, it is good to meet you.",
        "reasoning_content": "The user wants a greeting in one sentence. Keep it short and friendly."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 1184,
    "completion_tokens": 46,
    "total_tokens": 1230
  }
}

请求头

请求头

请求头 值 说明
Authorization Bearer YOUR_API_KEY 你的 API 密钥。在每个端点上,也可以用 x-api-key: YOUR_API_KEY 代替。
Content-Type application/json 必填。其他任何值都会返回 415。
x-request-id 可选。你自己的请求 id。它会原样出现在回复中。

回复头

请求头 说明
x-request-id 每个回复都有,包括错误和流:你发送的值;若未发送,则是 12 位十六进制字符。报告问题时请附上它。
content-type application/json,当 stream 为 true 时为 text/event-stream。

请求字段

只有 messages 是必需的。适用模型 一列列出了字段会改变回复的模型。托管开源权重模型即模型列表中的十二个 id;Shannon 3 系列为 shannon-3、shannon-3-pro、shannon-3.1 和 shannon-3.1-pro。 模型与定价

字段 类型 默认值 说明 适用模型
model string shannon-1.6-lite 作答的模型:模型列表中的 id。每个请求都要发送。匹配不区分大小写。未发布的 id 会返回 400 unknown model。 所有模型
messages array 必填。 对话内容,按从旧到新排列。参见下文的消息。 所有模型
stream boolean false true 会在回复生成的同时以 server-sent events 形式发送。 所有模型
max_tokens integer 4096 回复的上限,以 tokens 计。1 至 65,536 范围之外的值会被调整到该范围内。它同时也是请求运行期间从你的余额中预留的数量。参见下文的输出长度。 托管开源权重模型、shannon-1.6-lite、shannon-1.6-pro、shannon-coder-1
max_completion_tokens integer 与 max_tokens 相同。两者都发送时,使用 max_tokens。 托管开源权重模型、shannon-1.6-lite、shannon-1.6-pro、shannon-coder-1
temperature number 采样温度。在托管开源权重模型上默认值为 1,取值保持在 0 到 2 之间。 托管开源权重模型、shannon-1.6-lite、shannon-1.6-pro、shannon-coder-1
top_p number 0.95 核采样。取值保持在 0 到 1 之间。 托管开源权重模型
seed integer 采样器的种子,任意整数。若不提供,种子由模型和对话推导而来,因此同一请求发送两次会使用相同的种子。 托管开源权重模型
stop string | array 字符串或字符串数组。最多使用 4 个。答案会在第一个出现的停止文本之前结束;停止文本本身不会返回。 托管开源权重模型
reasoning_effort string high 模型回答前的推理量:off、low、medium 或 high。none 和 minimal 等同于 off,default 等同于 medium,max 等同于 high。其他任何值都会返回 400。 托管开源权重模型
reasoning object 同一设置的对象形式:{"effort": "low"}。两者都发送时,使用 reasoning_effort。 托管开源权重模型
tools array 模型可以调用的函数,每个形如 {"type": "function", "function": {"name", "description", "parameters"}}。模型的调用会在 tool_calls 中返回;由你的代码运行它们。 所有模型
tool_choice string | object auto "auto" 由模型决定。"required" 强制其调用工具。{"type": "function", "function": {"name": "…"}} 强制其调用指定工具。 托管开源权重模型
response_format object {"type": "json_object"} 返回 JSON 答案,{"type": "json_schema", "json_schema": {…}} 返回遵循你的 schema 的答案。 所有 Shannon 档位;托管开源权重模型按各 id 的列出情况而定
web_search boolean false true 让模型在回答前先搜索网页。 shannon-1.6-*、shannon-2-*、Shannon 3 系列

其他 OpenAI 字段,如 n、user、stream_options、parallel_tool_calls、presence_penalty、frequency_penalty、logit_bias、logprobs、metadata、store 和 prompt_cache_key,会被接受,以便现有客户端代码无需修改即可运行。它们不会改变回复:始终只有一个 choice,且流始终以用量结尾。

字段的 JSON 类型有误(例如 "max_tokens": "100")会返回 422。缺少 messages 的请求同样如此。

工具、结构化输出、推理和网页搜索各有专页: 函数调用, 结构化输出, 推理强度, 内置网页搜索.

带选项的请求

此请求设置了系统消息、采样字段和推理强度。它使用托管开源权重模型,这些设置都会生效。

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.shannon-ai.com/v1",
)

response = client.chat.completions.create(
    model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
    messages=[
        {"role": "system", "content": "You are a physics teacher. Answer in two sentences."},
        {"role": "user", "content": "Why is the sky blue?"},
    ],
    max_tokens=512,
    temperature=0.3,
    top_p=0.9,
    seed=7,
    stop=["\n\n"],
    reasoning_effort="low",
)

message = response.choices[0].message
print(message.reasoning_content)  # the reasoning
print(message.content)            # the answer
print(response.usage)

回复的结构与上文相同。在托管开源权重模型上,其 usage 多出两项细节:从缓存读取的 prompt tokens,以及用于推理的 tokens。

200 JSON
{
  "usage": {
    "prompt_tokens": 31,
    "completion_tokens": 62,
    "total_tokens": 93,
    "prompt_tokens_details": {
      "cached_tokens": 0
    },
    "completion_tokens_details": {
      "reasoning_tokens": 21
    }
  }
}

输出长度

max_tokens 有两个作用。第一,它是请求开始时从你的余额中预留的 tokens 数。回复完成后,该数量会被请求实际用掉的 tokens 取代。如果 max_tokens 大于你余额的剩余量,即使回复本身放得下,请求也会返回 429 Quota exceeded。发送较小的 max_tokens 可以少预留一些。

shannon-coder-1 在此端点上的计数方式不同:每个请求算作你的计划中的一次 Shannon Coder 调用,且不会为其预留 tokens。 限制与余额

第二,它限制这些模型上回复的长度:

模型 max_tokens 对回复长度所起的作用
shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 回复达到上限时停止。流随后以 finish_reason length 结束。
托管开源权重模型 答案文本在 max_tokens 处停止。推理不计入其中。小于 256 的值按 256 处理。

未提供 max_tokens 或 max_completion_tokens 时,值为 4,096。在 shannon-coder-1 上为 65,536。

消息

每条消息是带有 role 和 content 的对象。content 是字符串;当消息不止包含文本时,则是由多个部分组成的数组。

角色 说明 适用模型
system 给模型的指令。请放在最前面。在 Shannon 档位上,使用的是第一条 system 消息。 托管开源权重模型、shannon-1.6-*、shannon-2-*、shannon-coder-1
developer 按 system 读取。 托管开源权重模型
user 你的提问。在 Shannon 档位上,最后一条 user 消息是 prompt,其之前的消息是历史记录。 所有模型
assistant 模型先前的回复。在其后发送工具结果时,请保留它的 tool_calls。 所有模型
tool 工具调用的结果:tool_call_id 为该调用的 id,content 为字符串形式的结果。 所有模型

使用 Shannon 3 系列 id 时,请把必须遵守的指令放进 user 消息。

在 Shannon 档位上,既没有用户文本也没有 tools 的请求会返回 400 No user message provided。

内容部分

部分 说明 可用于
{"type": "text", "text": "…"} 纯文本。 所有模型
{"type": "image_url", "image_url": {"url": "…"}} 图片,以带 base64 内容的 data: URL 或 http(s) URL 形式提供。 Shannon 3 系列、shannon-1.6-lite、shannon-1.6-pro,以及列出图片输入的托管开源权重模型
{"type": "file", "source": {"type": "base64", "media_type": "application/pdf", "data": "…"}} 一个文档(PDF、Word、PowerPoint 或 Excel 文件),通过 base64 编码或 URL 提供。 Shannon 3 系列

大小、限制和完整的形式列表另有专页说明。 图片与文件

回复对象

字段 类型 说明
id string chatcmpl- 后接 32 位十六进制字符。
object string 始终为 chat.completion。
created integer 回复的时间,以 Unix 秒计。
model string 作答模型的规范 id。其拼写可能与你发送的 id 不同。
choices array 始终恰好一个 choice,其 index 为 0。
choices[0].message.role string 始终为 assistant。
choices[0].message.content string | null 答案文本。带有 tool_calls 时,在 Shannon 档位上为 null;托管开源权重模型可以在调用之外同时发送文本。
choices[0].message.reasoning_content string | null 模型在答案之前写下的推理内容;没有时为 null。
choices[0].message.tool_calls array 仅在模型调用工具时出现。每个条目包含 id、type(function)以及 function,其中 name 和 arguments 以 JSON 字符串形式给出。
choices[0].message.annotations array 仅出现在带 web_search: true 且搜索找到了结果的请求中。content 里的标记所指的每个来源对应一个 url_citation,包含 url、title、start_index 和 end_index(标记的位置,按字符计数,不含结束位置)。
choices[0].finish_reason string 回复结束的原因。见结束原因。
usage object 该请求的 tokens。见用量。
sources array 仅出现在带 web_search: true 且搜索找到了结果的请求中:交给模型的结果,每条包含 index、title 和 url。答案中的 [1] 就是 index 为 1 的条目。

结束原因

finish_reason 说明
stop 模型完成了回答,或出现了某个 stop 字符串。
tool_calls 模型调用了一个或多个工具。请运行它们,并在 tool 消息中发送结果。
length 回复在输出上限处被截断。会在 shannon-1.6-lite、shannon-1.6-pro、shannon-coder-1 和 Shannon 3 系列的流中报告。

非流式回复报告 stop 或 tool_calls。

用量

字段 类型 说明 可用于
usage.prompt_tokens integer 输入 tokens。 所有模型
usage.completion_tokens integer 输出 tokens:推理、答案和工具调用的总和。 所有模型
usage.total_tokens integer prompt_tokens 加 completion_tokens。 所有模型
usage.prompt_tokens_details.cached_tokens integer prompt_tokens 中从 prompt 缓存读取的部分。 托管开源权重模型
usage.completion_tokens_details.reasoning_tokens integer completion_tokens 中用于推理的部分。 托管开源权重模型

在托管开源权重模型上,prompt_tokens 是用该模型自己的分词器统计的消息和工具定义,加上任何图片的 tokens。tokens 计数端点会在你发送之前返回相同的数字。 Token 计数

在 Shannon 档位上,prompt_tokens 统计模型为写出回复所读取的全部内容,因此比你的消息文本本身要大。

流式传输

将 stream 设为 true 时,回复以 chat.completion.chunk 事件到达,并以 data: [DONE] 结束。在它之前的最后一个数据块带有 finish_reason 和 usage;无需 stream_options。数据块的结构、保活行以及流内的错误另有专页说明。 流式

错误

错误是带 error 成员的 JSON 对象。检查按以下顺序进行:API 密钥、请求体、模型 id,然后是余额。下表列出此端点最常返回的错误。完整列表及重试建议见专页。 错误处理

400 JSON
{
  "error": {
    "type": "invalid_request_error",
    "message": "unknown model: no-such-model"
  }
}
状态码 类型 消息 时机
401 authentication_error Missing authentication
Invalid API key
未发送 API 密钥,或密钥未知或已被撤销。
400 invalid_request_error unknown model: <id> model 不是已发布的 id。
400 invalid_request_error No user message provided Shannon 档位:请求既没有用户文本,也没有 tools。
400 invalid_request_error <id> does not accept image input 向不支持图片输入的托管开源权重模型发送了图片部分。
400 invalid_request_error <id> does not accept response_format 向不支持结构化输出的托管开源权重模型发送了 response_format。
400 invalid_request_error unknown reasoning effort '<value>'; expected off, low, medium or high reasoning_effort 的值不在列表内。
422 invalid_request_error Failed to deserialize the JSON body into the target type: … 缺少 messages,或某个字段的 JSON 类型有误。
429 rate_limit_error Quota exceeded. Upgrade your plan at shannon-ai.com/plan max_tokens 大于你余额的剩余量。
429 rate_limit_error Too many requests. Retry in <n>s. 限流保护:你的账户在一分钟内发送了超过 120 个请求。
500 server_error The model backend failed to answer. Please retry. 模型未生成回复。请重新发送请求。
502 api_error The model backend failed to answer. Please retry. 同上,适用于 Shannon 3 系列和托管开源权重模型。