Token 计数
在发送之前,计算一段文本或整个请求的 tokens 数。
POST https://api.shannon-ai.com/v1/tokenize
POST https://api.shannon-ai.com/v1/messages/count_tokens
两个端点都使用你所指定模型的分词器计数,不会运行任何模型。它们覆盖托管开源权重模型。/v1/tokenize 接受纯文本或 Chat Completions 对话。/v1/messages/count_tokens 接受 Anthropic Messages 格式的请求,Anthropic SDK 和 Claude Code 发出的正是这种调用。
计数是免费的。调用需要你的 API 密钥,不会从你的余额中扣除任何内容,也不会出现在你的用量记录中。
为文本计数
发送 model 和 text。文本按原样计数,周围没有聊天格式。
import requests
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world",
},
)
print(response.json()["tokens"]) const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
text: "Hello, world",
}),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world"
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 3
} 本页回复中的数字只是示例。同一段文本在不同模型上的计数不同。
为聊天请求计数
发送 model 和 messages,如果请求带有工具,再加上 tools,与你发送给 /v1/chat/completions 的方式完全相同。回复是整个输入的大小。
import requests
request = {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"},
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
],
}
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json=request,
)
print(response.json()["tokens"]) const request = {
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
messages: [
{ role: "system", content: "You are a concise assistant." },
{ role: "user", content: "What is the weather in Paris?" },
],
tools: [
{
type: "function",
function: {
name: "get_weather",
description: "Current weather for a city",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
},
},
},
],
};
const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify(request),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}
]
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 164
} /v1/tokenize 的字段
| 字段 | 类型 | 说明 |
|---|---|---|
model | string | 必填。 托管开源权重模型的 id。大小写视为相同。 |
text | string | 按原样计数的文本,没有聊天格式。最多 4,000,000 字节。发送 text 或 messages;两者都有时,按 text 计数。 |
messages | array | Chat Completions 格式的聊天消息。它们按请求的完整输入计数:每条消息连同模型的聊天模板围绕它添加的格式。 |
tools | array | 要计入的工具定义。与 messages 一起使用。 |
回复是一个 JSON 对象,包含以下字段:
| 字段 | 类型 | 说明 |
|---|---|---|
model | string | 此次计数所针对的模型 id,采用其已发布的写法。 |
tokens | integer | 使用 text 时:该文本的 tokens。使用 messages 时:整个输入的 tokens,包括图片。 |
为 Messages 请求计数
发送你会发送给 /v1/messages 的请求体:model、messages,以及使用时的 system 和 tools。官方 Anthropic SDK 通过 messages.count_tokens 调用此端点。
import anthropic
client = anthropic.Anthropic(
api_key="YOUR_API_KEY",
base_url="https://api.shannon-ai.com",
)
count = client.messages.count_tokens(
model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system="You are a concise assistant.",
messages=[
{"role": "user", "content": "Summarise the attached report."}
],
)
print(count.input_tokens) import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
apiKey: "YOUR_API_KEY",
baseURL: "https://api.shannon-ai.com",
});
const count = await client.messages.countTokens({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system: "You are a concise assistant.",
messages: [
{ role: "user", content: "Summarise the attached report." },
],
});
console.log(count.input_tokens); curl https://api.shannon-ai.com/v1/messages/count_tokens \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"system": "You are a concise assistant.",
"messages": [
{"role": "user", "content": "Summarise the attached report."}
]
}' {
"input_tokens": 21
} /v1/messages/count_tokens 请求体中可以使用的各个字段
| 字段 | 类型 | 说明 |
|---|---|---|
model | string | 必填。 托管开源权重模型的 id。 |
messages | array | 必填。 Anthropic Messages 格式的消息。text、image、tool_use 和 tool_result 块都会被计数。 |
system | string | array | 系统提示词:字符串或文本块数组。 |
tools | array | 带有 name、description 和 input_schema 的工具定义。 |
为兼容性而接受,对计数没有影响:tool_choice, max_tokens, temperature, top_p, stop_sequences, stream, thinking。你可以原样传入真实请求的请求体。
回复是一个 JSON 对象,包含以下字段:
| 字段 | 类型 | 说明 |
|---|---|---|
input_tokens | integer | 整个输入的 tokens:系统提示词、消息、工具和图片。 |
支持的模型
两个端点都为托管开源权重模型计数。GET /v1/models 会在每个支持它们的模型的 endpoints 中列出 /v1/tokenize 和 /v1/messages/count_tokens。其他任何 model 值,包括 Shannon 的 id,都会返回 400。
DeepSeek-V4-Pro-0813-3BIT-REAPGLM-5.2-3BIT-REAPKimi-K3-3BIT-REAPNemotron3Ultra-3BIT-REAPMiniMax-M3-3BIT-REAPDeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAPKimi-K2.6-W4A16-AUTOROUND-REAPLaguna-S-2.1-W4A16-AUTOROUND-REAPinkling-W4A16-AUTOROUND-REAPMiMo-V2.5-Pro-W8A16MiMo-V2.5-W8A16Hy3-W8A16
对于 Shannon 模型,请从回复的 usage 对象中读取 token 数。
计数的方式
每个模型都用它自己的分词器和聊天模板来计数。不使用按字符或单词得出的估算。
| 计数内容 | 规则 |
|---|---|
| 文本 | 按发送的原样计算字符串的 tokens。空字符串计为 0。 |
| 消息 | 消息和工具按模型自己的聊天模板排版,直到回复开始之处,整个 prompt 都会被计数。 |
| 角色 | system、user、assistant 和 tool 消息都会被计数。developer 按 system 计。既没有内容也没有工具调用的消息不增加计数。 |
| 工具调用与结果 | 在两个端点上,之前助手轮次的工具调用及其结果都是计数的一部分。 |
| 图片 | 在请求体内发送的图片(base64 或 data: URL)每个 28 × 28 像素的图块增加一个 token:ceil(width / 28) × ceil(height / 28)。以 http(s) URL 给出的图片,这些端点不会下载,按 1,024 计。 |
示例:一张 1,024 × 768 像素的图片计为 ceil(1024 / 28) × ceil(768 / 28) = 37 × 28 = 1,036 tokens。
计数以及请求的计费方式
整个请求的计数方式,与使用相同模型、消息和工具的真实请求的输入计数相同。回复把该数字报告为:Chat Completions 上的 usage.prompt_tokens,Responses 上的 usage.input_tokens,以及 Messages 上的 usage.input_tokens 加 usage.cache_read_input_tokens。
- 该计数是应用缓存输入折扣之前的输入。真实请求可能从缓存中读取其中一部分输入,并按缓存价格对该部分计费。 Prompt 缓存
- 以
http(s)URL 给出的图片在这里按 1,024 计。真实请求会下载图片,并根据其像素尺寸计数,因此两个数字可能不同。以 base64 发送图片即可得到相同的数字。 - 输出不属于计数。真实请求的回复会另外按输出 tokens 计费,包括推理。
text计数不含聊天格式。用它来衡量文档或 prompt 的一部分,用messages形式来衡量请求。
要把计数换算为费用,请将其乘以该模型每 1M tokens 的输入价格。 模型与定价
限制
| 限制 | 值 | 超出时 |
|---|---|---|
text 的长度 | 4,000,000 字节(UTF-8) | 413,消息为 text too long |
| 请求体 | 32 MiB | 413 |
| 每个请求 | 一段文本或一个对话 | 要为多段文本计数,请为每段文本发送一个请求。 |
计数调用不计入每分钟 120 个请求的限额。 限制与余额
错误
| 状态 | 类型 | 消息 | 何时 |
|---|---|---|---|
400 | invalid_request_error | tokenize is available for the hosted open models; unknown model: <model> | /v1/tokenize 的 model 不是托管开源权重模型的 id。 |
400 | invalid_request_error | count_tokens is available for the hosted open models; unknown model: <model> | /v1/messages/count_tokens 的 model 不是托管开源权重模型的 id,或没有 model。 |
400 | invalid_request_error | send `text` or `messages` | /v1/tokenize 既没有 text 也没有 messages。 |
401 | authentication_error | Missing authentication / Invalid API key | 没有发送密钥,或密钥无效。 |
413 | invalid_request_error | text too long | text 超过 4,000,000 字节。超过 32 MiB 的请求体同样会返回 413。 |
415 | invalid_request_error | Expected request with `Content-Type: application/json` | 请求没有 JSON 内容类型。 |
422 | invalid_request_error | Failed to deserialize the JSON body into the target type: … | 缺少必填字段(/v1/tokenize 上的 model,/v1/messages/count_tokens 上的 messages),或某个字段的类型错误。 |
503 | api_error | token counting is temporarily unavailable for this model | 目前无法对该模型进行计数。请稍后重试。 |
/v1/tokenize 以 OpenAI 的结构返回错误。在 /v1/messages/count_tokens 上,端点自身的错误(针对模型的 400、503)采用 Anthropic 的结构,而 401、413、415 和 422 采用 OpenAI 的结构。请先读取状态码,然后读取两种结构中都有的 error.type 和 error.message。
{
"error": {
"type": "invalid_request_error",
"message": "tokenize is available for the hosted open models; unknown model: shannon-3"
}
} {
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "count_tokens is available for the hosted open models; unknown model: shannon-3"
}
}