トークンカウント
テキストまたはリクエスト全体のトークンを、送信前に数えます。
POST https://api.shannon-ai.com/v1/tokenize
POST https://api.shannon-ai.com/v1/messages/count_tokens
どちらのエンドポイントも、指定したモデルのトークナイザーで数え、モデルは実行されません。対象はホスト型オープンウェイトモデルです。/v1/tokenize はプレーンテキスト、または Chat Completions の会話を受け付けます。/v1/messages/count_tokens は Anthropic Messages 形式のリクエストを受け付けます。これは Anthropic SDK や Claude Code が行う呼び出しです。
カウントは無料です。呼び出しには API キーが必要ですが、残高からは何も引かれず、使用量ログにも表示されません。
テキストのトークンを数える
model と text を送信します。テキストは、チャット用の整形なしで、そのまま数えられます。
import requests
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world",
},
)
print(response.json()["tokens"]) const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
text: "Hello, world",
}),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world"
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 3
} このページの応答に含まれる数値は例です。同じテキストでも、モデルが異なればカウントも異なります。
チャットリクエストのトークンを数える
model と messages を、リクエストにツールがある場合は tools も、/v1/chat/completions に送るのとまったく同じ形で送信します。応答は、入力全体のサイズです。
import requests
request = {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"},
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
],
}
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json=request,
)
print(response.json()["tokens"]) const request = {
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
messages: [
{ role: "system", content: "You are a concise assistant." },
{ role: "user", content: "What is the weather in Paris?" },
],
tools: [
{
type: "function",
function: {
name: "get_weather",
description: "Current weather for a city",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
},
},
},
],
};
const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify(request),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}
]
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 164
} /v1/tokenize に指定できるフィールドの一覧
| フィールド | 型 | 説明 |
|---|---|---|
model | string | 必須。 ホスト型オープンウェイトモデルの ID。大文字と小文字は同じに扱われます。 |
text | string | チャット用の整形なしで、そのまま数えるテキスト。最大 4,000,000 バイト。text または messages を送信してください。両方ある場合は text が数えられます。 |
messages | array | Chat Completions 形式のチャットメッセージ。リクエストの入力全体として数えられます。すべてのメッセージに、モデルのチャットテンプレートが周囲に付ける整形が含まれます。 |
tools | array | カウントに含めるツール定義。messages と一緒に使います。 |
応答は、次のフィールドを持つ JSON オブジェクトです。
| フィールド | 型 | 説明 |
|---|---|---|
model | string | カウントの対象となったモデル ID(公開されている表記)。 |
tokens | integer | text の場合:テキストのトークン。messages の場合:画像を含む入力全体のトークン。 |
Messages リクエストのトークンを数える
/v1/messages に送信するのと同じボディを送信します:model、messages、そして使う場合は system と tools。公式の Anthropic SDK は、messages.count_tokens を通じてこのエンドポイントを呼び出します。
import anthropic
client = anthropic.Anthropic(
api_key="YOUR_API_KEY",
base_url="https://api.shannon-ai.com",
)
count = client.messages.count_tokens(
model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system="You are a concise assistant.",
messages=[
{"role": "user", "content": "Summarise the attached report."}
],
)
print(count.input_tokens) import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
apiKey: "YOUR_API_KEY",
baseURL: "https://api.shannon-ai.com",
});
const count = await client.messages.countTokens({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system: "You are a concise assistant.",
messages: [
{ role: "user", content: "Summarise the attached report." },
],
});
console.log(count.input_tokens); curl https://api.shannon-ai.com/v1/messages/count_tokens \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"system": "You are a concise assistant.",
"messages": [
{"role": "user", "content": "Summarise the attached report."}
]
}' {
"input_tokens": 21
} /v1/messages/count_tokens に指定できるフィールドの一覧
| フィールド | 型 | 説明 |
|---|---|---|
model | string | 必須。 ホスト型オープンウェイトモデルの ID。 |
messages | array | 必須。 Anthropic Messages 形式のメッセージ。text、image、tool_use、tool_result ブロックが数えられます。 |
system | string | array | システムプロンプト:文字列、またはテキストブロックの配列。 |
tools | array | name、description、input_schema を持つツール定義。 |
互換性のために受け付けられ、カウントには影響しません:tool_choice, max_tokens, temperature, top_p, stop_sequences, stream, thinking。実際のリクエストのボディをそのまま渡せます。
応答は、次のフィールドを持つ JSON オブジェクトです。
| フィールド | 型 | 説明 |
|---|---|---|
input_tokens | integer | 入力全体(システムプロンプト、メッセージ、ツール、画像)のトークン。 |
対応モデル
どちらのエンドポイントも、ホスト型オープンウェイトモデルを対象に数えます。GET /v1/models は、対応する各モデルの endpoints に /v1/tokenize と /v1/messages/count_tokens を一覧表示します。それ以外の model の値(Shannon の ID を含む)には、400 で応答します。
DeepSeek-V4-Pro-0813-3BIT-REAPGLM-5.2-3BIT-REAPKimi-K3-3BIT-REAPNemotron3Ultra-3BIT-REAPMiniMax-M3-3BIT-REAPDeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAPKimi-K2.6-W4A16-AUTOROUND-REAPLaguna-S-2.1-W4A16-AUTOROUND-REAPinkling-W4A16-AUTOROUND-REAPMiMo-V2.5-Pro-W8A16MiMo-V2.5-W8A16Hy3-W8A16
Shannon モデルでは、トークン数を応答の usage オブジェクトから読み取ってください。
カウントの方法
各モデルは、そのモデル独自のトークナイザーとチャットテンプレートで数えられます。文字数や単語数からの推定は使いません。
| カウント対象 | ルール |
|---|---|
| テキスト | 送信された文字列のトークン。空の文字列は 0 と数えられます。 |
| メッセージ | メッセージとツールは、応答が始まる直前まで、モデル独自のチャットテンプレートで組み立てられ、そのプロンプト全体が数えられます。 |
| ロール | system、user、assistant、tool のメッセージが数えられます。developer は system として数えられます。内容もツール呼び出しもないメッセージは、何も加算されません。 |
| ツール呼び出しと結果 | それ以前のアシスタントターンのツール呼び出しと、その結果は、どちらのエンドポイントでもカウントに含まれます。 |
| 画像 | 本文内で送信された画像(base64 または data: URL)は、28 × 28 ピクセルのパッチごとに一トークンが加算されます:ceil(width / 28) × ceil(height / 28)。http(s) URL で指定された画像は、これらのエンドポイントではダウンロードされず、1,024 として数えられます。 |
例:1,024 × 768 ピクセルの画像は、ceil(1024 / 28) × ceil(768 / 28) = 37 × 28 = 1,036 トークンと数えられます。
カウントとリクエストの課金額
リクエスト全体のカウントは、同じモデル、メッセージ、ツールを使った実際のリクエストの入力カウントと同じ方法で行われます。応答はその数値を、Chat Completions では usage.prompt_tokens、Responses では usage.input_tokens、Messages では usage.input_tokens と usage.cache_read_input_tokens の合計として報告します。
- カウントは、キャッシュ入力の割引前の入力です。実際のリクエストでは、その入力の一部がキャッシュから読み取られ、その部分がキャッシュレートで課金されることがあります。 プロンプトキャッシュ
http(s)URL で指定された画像は、ここでは 1,024 として数えられます。実際のリクエストは画像をダウンロードし、ピクセル単位のサイズから数えるため、二つの数値は異なることがあります。同じ数値を得るには、画像を base64 で送信してください。- 出力はカウントに含まれません。実際のリクエストの応答は、推論を含めて、これに加えて出力トークンとして課金されます。
textのカウントには、チャット用の整形が含まれません。ドキュメントやプロンプトの一部を測るのに使い、リクエストを測るにはmessages形式を使ってください。
カウントをコストに換算するには、モデルの 1M トークンあたりの入力価格を掛けます。 モデルと料金
上限
| 上限 | 値 | 上限を超えた場合 |
|---|---|---|
text の長さ | 4,000,000 バイト(UTF-8) | メッセージ text too long を伴う 413 |
| リクエストボディ | 32 MiB | 413 |
| リクエストごと | ひとつのテキスト、またはひとつの会話 | 複数のテキストを数えるには、テキストごとにリクエストを送信してください。 |
カウント呼び出しは、一分間 120 リクエストの上限にはカウントされません。 上限と残高
エラー
| ステータス | 型 | メッセージ | 発生条件 |
|---|---|---|---|
400 | invalid_request_error | tokenize is available for the hosted open models; unknown model: <model> | ホスト型オープンウェイトの ID ではない model を指定した /v1/tokenize。 |
400 | invalid_request_error | count_tokens is available for the hosted open models; unknown model: <model> | ホスト型オープンウェイトの ID ではない model を指定した、または model を指定しない /v1/messages/count_tokens。 |
400 | invalid_request_error | send `text` or `messages` | text も messages も指定しない /v1/tokenize。 |
401 | authentication_error | Missing authentication / Invalid API key | キーが送信されていないか、キーが有効ではありません。 |
413 | invalid_request_error | text too long | text が 4,000,000 バイトを超えています。32 MiB を超えるボディにも 413 で応答します。 |
415 | invalid_request_error | Expected request with `Content-Type: application/json` | リクエストに JSON のコンテンツタイプがありません。 |
422 | invalid_request_error | Failed to deserialize the JSON body into the target type: … | 必須フィールド(/v1/tokenize では model、/v1/messages/count_tokens では messages)がないか、フィールドの型が正しくありません。 |
503 | api_error | token counting is temporarily unavailable for this model | 現在、このモデルではカウントできません。後でもう一度お試しください。 |
/v1/tokenize は OpenAI 形式でエラーを返します。/v1/messages/count_tokens では、このエンドポイント自体のエラー(モデルに関する 400、503)は Anthropic 形式で、401、413、415、422 は OpenAI 形式で届きます。まずステータスコードを読み、次に、どちらの形式にもある error.type と error.message を読んでください。
{
"error": {
"type": "invalid_request_error",
"message": "tokenize is available for the hosted open models; unknown model: shannon-3"
}
} {
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "count_tokens is available for the hosted open models; unknown model: shannon-3"
}
}