Pagbibilang ng token
Bilangin ang tokens ng isang text o ng buong request bago mo ito ipadala.
POST https://api.shannon-ai.com/v1/tokenize
POST https://api.shannon-ai.com/v1/messages/count_tokens
Parehong bumibilang ang dalawang endpoint gamit ang tokenizer ng model na pinangalanan mo, at walang model na tumatakbo. Sakop nila ang mga hosted open-weight model. Ang /v1/tokenize ay tumatanggap ng payak na text o Chat Completions conversation. Ang /v1/messages/count_tokens ay tumatanggap ng request sa Anthropic Messages format, na siyang call na ginagawa ng Anthropic SDK at Claude Code.
Libre ang pagbibilang. Kailangan ng iyong API key ang isang call, walang kinukuha sa iyong balance at hindi ito lumalabas sa iyong usage log.
Bilangin ang isang text
Ipadala ang model at text. Binibilang ang text ayon sa kung ano ito, na walang chat formatting sa paligid.
import requests
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world",
},
)
print(response.json()["tokens"]) const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
text: "Hello, world",
}),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world"
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 3
} Mga halimbawa ang mga numero sa mga reply sa pahinang ito. Ang parehong text ay nagbibigay ng ibang bilang sa ibang model.
Bilangin ang chat request
Ipadala ang model at messages, kasama ang tools kapag mayroon ang request, eksakto kung paano mo ipapadala ang mga ito sa /v1/chat/completions. Ang reply ay ang laki ng buong input.
import requests
request = {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"},
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
],
}
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json=request,
)
print(response.json()["tokens"]) const request = {
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
messages: [
{ role: "system", content: "You are a concise assistant." },
{ role: "user", content: "What is the weather in Paris?" },
],
tools: [
{
type: "function",
function: {
name: "get_weather",
description: "Current weather for a city",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
},
},
},
],
};
const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify(request),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}
]
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 164
} Mga field ng /v1/tokenize
| Field | Type | Paglalarawan |
|---|---|---|
model | string | Kinakailangan. Isang hosted open-weight model id. Pareho ang pagtrato sa malaki at maliit na titik. |
text | string | Isang text na bibilangin ayon sa kung ano ito, na walang chat formatting. Hanggang 4,000,000 byte. Ipadala ang text o messages; kapag parehong naroroon, ang text ang binibilang. |
messages | array | Mga chat message sa Chat Completions format. Binibilang ang mga ito bilang buong input ng isang request: bawat message kasama ang formatting na inilalagay ng chat template ng model sa paligid nito. |
tools | array | Mga tool definition na isasama sa bilang. Ginagamit kasama ng messages. |
Ang reply ay isang JSON object na may mga field na ito:
| Field | Type | Paglalarawan |
|---|---|---|
model | string | Ang model id kung saan ginawa ang bilang, sa inilathalang baybay nito. |
tokens | integer | Sa text: ang mga token ng text. Sa messages: ang mga token ng buong input, kasama ang mga image. |
Bilangin ang Messages request
Ipadala ang body na ipapadala mo sa /v1/messages: model, messages, at system at tools kapag ginagamit mo ang mga ito. Tinatawag ng mga opisyal na Anthropic SDK ang endpoint na ito sa pamamagitan ng messages.count_tokens.
import anthropic
client = anthropic.Anthropic(
api_key="YOUR_API_KEY",
base_url="https://api.shannon-ai.com",
)
count = client.messages.count_tokens(
model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system="You are a concise assistant.",
messages=[
{"role": "user", "content": "Summarise the attached report."}
],
)
print(count.input_tokens) import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
apiKey: "YOUR_API_KEY",
baseURL: "https://api.shannon-ai.com",
});
const count = await client.messages.countTokens({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system: "You are a concise assistant.",
messages: [
{ role: "user", content: "Summarise the attached report." },
],
});
console.log(count.input_tokens); curl https://api.shannon-ai.com/v1/messages/count_tokens \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"system": "You are a concise assistant.",
"messages": [
{"role": "user", "content": "Summarise the attached report."}
]
}' {
"input_tokens": 21
} Mga field ng /v1/messages/count_tokens
| Field | Type | Paglalarawan |
|---|---|---|
model | string | Kinakailangan. Isang hosted open-weight model id. |
messages | array | Kinakailangan. Mga message sa Anthropic Messages format. Binibilang ang mga block na text, image, tool_use at tool_result. |
system | string | array | Ang system prompt: isang string o array ng mga text block. |
tools | array | Mga tool definition na may name, description at input_schema. |
Tinatanggap para sa compatibility, na walang epekto sa bilang: tool_choice, max_tokens, temperature, top_p, stop_sequences, stream, thinking. Maaari mong ipasa nang walang pagbabago ang body ng tunay na request.
Ang reply ay isang JSON object na may mga field na ito:
| Field | Type | Paglalarawan |
|---|---|---|
input_tokens | integer | Ang mga token ng buong input: system prompt, messages, tools at mga image. |
Mga sinusuportahang model
Parehong bumibilang ang dalawang endpoint para sa mga hosted open-weight model. Inililista ng GET /v1/models ang /v1/tokenize at /v1/messages/count_tokens sa endpoints ng bawat model na sumusuporta sa mga ito. Ang anumang ibang value ng model, kasama ang mga Shannon id, ay sinasagot ng 400.
DeepSeek-V4-Pro-0813-3BIT-REAPGLM-5.2-3BIT-REAPKimi-K3-3BIT-REAPNemotron3Ultra-3BIT-REAPMiniMax-M3-3BIT-REAPDeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAPKimi-K2.6-W4A16-AUTOROUND-REAPLaguna-S-2.1-W4A16-AUTOROUND-REAPinkling-W4A16-AUTOROUND-REAPMiMo-V2.5-Pro-W8A16MiMo-V2.5-W8A16Hy3-W8A16
Para sa Shannon model, basahin ang bilang ng token mula sa usage object ng isang reply.
Paano ginagawa ang bilang
Binibilang ang bawat model gamit ang sarili nitong tokenizer at sarili nitong chat template. Walang tantiyang mula sa mga character o salita ang ginagamit.
| Ano ang binibilang | Patakaran |
|---|---|
| Isang text | Ang mga token ng string ayon sa pagkakapadala. Ang walang lamang string ay binibilang na 0. |
| Mga message | Inilalatag ang mga message at tool gamit ang sariling chat template ng model, hanggang sa puntong magsisimula ang reply, at binibilang ang buong prompt na iyon. |
| Mga role | Binibilang ang mga message na system, user, assistant at tool. Ang developer ay binibilang bilang system. Ang message na walang content at walang tool call ay walang idinaragdag. |
| Mga tool call at resulta | Bahagi ng bilang ang mga tool call ng mga naunang assistant turn at ang mga resulta ng mga ito, sa parehong endpoint. |
| Mga image | Ang image na ipinadala sa loob ng body (base64 o data: URL) ay nagdaragdag ng isang token kada 28 × 28 pixel patch: ceil(width / 28) × ceil(height / 28). Ang image na ibinigay bilang http(s) URL ay hindi dina-download ng mga endpoint na ito at binibilang na 1,024. |
Halimbawa: ang image na 1,024 × 768 pixel ay binibilang na ceil(1024 / 28) × ceil(768 / 28) = 37 × 28 = 1,036 token.
Ang bilang at kung ano ang sinisingil sa request
Ginagawa ang bilang ng buong request sa parehong paraan ng input count ng tunay na request na may parehong model, messages at tools. Iniuulat ng reply ang numerong iyon bilang usage.prompt_tokens sa Chat Completions, bilang usage.input_tokens sa Responses, at bilang usage.input_tokens dagdag ang usage.cache_read_input_tokens sa Messages.
- Ang bilang ay ang input bago ang cached-input discount. Ang tunay na request ay maaaring magbasa ng bahagi ng input na iyon mula sa cache at singilin ang bahaging iyon sa cached rate. Prompt caching
- Ang image na ibinigay bilang
http(s)URL ay binibilang na 1,024 dito. Dina-download ng tunay na request ang image at binibilang ito mula sa laki nito sa pixel, kaya maaaring magkaiba ang dalawang numero. Ipadala ang image bilang base64 para makuha ang parehong numero. - Hindi bahagi ng bilang ang output. Ang reply ng tunay na request ay sinisingil bilang output tokens bukod pa rito, kasama ang reasoning.
- Walang chat formatting ang bilang ng
text. Gamitin ito para sukatin ang dokumento o bahagi ng prompt, at ang anyongmessagespara sukatin ang request.
Para gawing gastos ang bilang, i-multiply ito sa input price ng model kada 1M tokens. Mga model at presyo
Mga limitasyon
| Limitasyon | Value | Lampas dito |
|---|---|---|
Haba ng text | 4,000,000 byte (UTF-8) | 413 na may mensaheng text too long |
| Request body | 32 MiB | 413 |
| Kada request | Isang text o isang conversation | Magpadala ng isang request kada text para magbilang ng ilang text. |
Hindi binibilang ang mga counting call sa limitasyong 120 request kada minuto. Mga limitasyon at balance
Mga error
| Status | Type | Mensahe | Kailan |
|---|---|---|---|
400 | invalid_request_error | tokenize is available for the hosted open models; unknown model: <model> | Ang /v1/tokenize na may model na hindi hosted open-weight id. |
400 | invalid_request_error | count_tokens is available for the hosted open models; unknown model: <model> | Ang /v1/messages/count_tokens na may model na hindi hosted open-weight id, o walang model. |
400 | invalid_request_error | send `text` or `messages` | Ang /v1/tokenize na walang text at walang messages. |
401 | authentication_error | Missing authentication / Invalid API key | Walang key na naipadala, o hindi valid ang key. |
413 | invalid_request_error | text too long | Mas mahaba sa 4,000,000 byte ang text. Ang body na lampas sa 32 MiB ay sinasagot din ng 413. |
415 | invalid_request_error | Expected request with `Content-Type: application/json` | Walang JSON content type ang request. |
422 | invalid_request_error | Failed to deserialize the JSON body into the target type: … | Nawawala ang isang kinakailangang field (model sa /v1/tokenize, messages sa /v1/messages/count_tokens) o mali ang type ng isang field. |
503 | api_error | token counting is temporarily unavailable for this model | Hindi magawa ang bilang para sa model na ito sa ngayon. Subukan muli mamaya. |
Nagbabalik ang /v1/tokenize ng mga error sa hugis ng OpenAI. Sa /v1/messages/count_tokens, ang mga error ng mismong endpoint (400 para sa model, 503) ay dumarating sa hugis ng Anthropic, at ang 401, 413, 415 at 422 ay dumarating sa hugis ng OpenAI. Basahin muna ang status code, saka ang error.type at error.message, na naroroon sa parehong hugis.
{
"error": {
"type": "invalid_request_error",
"message": "tokenize is available for the hosted open models; unknown model: shannon-3"
}
} {
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "count_tokens is available for the hosted open models; unknown model: shannon-3"
}
}