การนับ token
นับ tokens ของข้อความหรือของทั้งคำขอก่อนที่คุณจะส่ง
POST https://api.shannon-ai.com/v1/tokenize
POST https://api.shannon-ai.com/v1/messages/count_tokens
ทั้งสอง endpoint นับด้วย tokenizer ของโมเดลที่คุณระบุ และไม่มีโมเดลทำงาน ครอบคลุมโมเดล open-weight ที่โฮสต์ /v1/tokenize รับข้อความธรรมดาหรือบทสนทนา Chat Completions /v1/messages/count_tokens รับคำขอในรูปแบบ Anthropic Messages ซึ่งเป็นการเรียกที่ Anthropic SDK และ Claude Code ทำ
การนับไม่มีค่าใช้จ่าย การเรียกต้องใช้ API key ของคุณ ไม่หักจากยอดคงเหลือ และไม่ปรากฏในบันทึกการใช้งานของคุณ
นับข้อความ
ส่ง model และ text ข้อความจะถูกนับตามที่เป็น โดยไม่มีการจัดรูปแบบแชตล้อมรอบ
import requests
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world",
},
)
print(response.json()["tokens"]) const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
text: "Hello, world",
}),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world"
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 3
} ตัวเลขในคำตอบของหน้านี้เป็นตัวอย่าง ข้อความเดียวกันให้จำนวนที่นับต่างกันบนโมเดลที่ต่างกัน
นับคำขอแชต
ส่ง model และ messages พร้อม tools เมื่อคำขอมี เหมือนที่จะส่งไปยัง /v1/chat/completions ทุกประการ คำตอบคือขนาดของอินพุตทั้งหมด
import requests
request = {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"},
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
],
}
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json=request,
)
print(response.json()["tokens"]) const request = {
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
messages: [
{ role: "system", content: "You are a concise assistant." },
{ role: "user", content: "What is the weather in Paris?" },
],
tools: [
{
type: "function",
function: {
name: "get_weather",
description: "Current weather for a city",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
},
},
},
],
};
const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify(request),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}
]
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 164
} ฟิลด์ของ /v1/tokenize
| ฟิลด์ | ประเภท | คำอธิบาย |
|---|---|---|
model | string | จำเป็น model id ของโมเดล open-weight ที่โฮสต์ ตัวพิมพ์ใหญ่และเล็กถือเป็นอย่างเดียวกัน |
text | string | ข้อความที่นับตามที่เป็น โดยไม่มีการจัดรูปแบบแชต ไม่เกิน 4,000,000 ไบต์ ส่ง text หรือ messages เมื่อมีทั้งสอง จะนับ text |
messages | array | ข้อความแชตในรูปแบบ Chat Completions ถูกนับเป็นอินพุตเต็มของคำขอ: ทุกข้อความพร้อมการจัดรูปแบบที่ chat template ของโมเดลใส่รอบข้อความนั้น |
tools | array | นิยาม tool ที่ต้องรวมในการนับ ใช้ร่วมกับ messages |
คำตอบคือออบเจ็กต์ JSON ที่มีฟิลด์เหล่านี้:
| ฟิลด์ | ประเภท | คำอธิบาย |
|---|---|---|
model | string | model id ที่ใช้นับ ในรูปที่เผยแพร่ |
tokens | integer | เมื่อใช้ text: tokens ของข้อความ เมื่อใช้ messages: tokens ของอินพุตทั้งหมด รวมรูปภาพ |
นับคำขอ Messages
ส่ง body เดียวกับที่ส่งไป /v1/messages: model, messages และ system กับ tools เมื่อใช้ Anthropic SDK อย่างเป็นทางการเรียก endpoint นี้ผ่าน messages.count_tokens
import anthropic
client = anthropic.Anthropic(
api_key="YOUR_API_KEY",
base_url="https://api.shannon-ai.com",
)
count = client.messages.count_tokens(
model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system="You are a concise assistant.",
messages=[
{"role": "user", "content": "Summarise the attached report."}
],
)
print(count.input_tokens) import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
apiKey: "YOUR_API_KEY",
baseURL: "https://api.shannon-ai.com",
});
const count = await client.messages.countTokens({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system: "You are a concise assistant.",
messages: [
{ role: "user", content: "Summarise the attached report." },
],
});
console.log(count.input_tokens); curl https://api.shannon-ai.com/v1/messages/count_tokens \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"system": "You are a concise assistant.",
"messages": [
{"role": "user", "content": "Summarise the attached report."}
]
}' {
"input_tokens": 21
} ฟิลด์ที่ใช้กับ endpoint สำหรับนับ token แบบ Messages (/v1/messages/count_tokens)
| ฟิลด์ | ประเภท | คำอธิบาย |
|---|---|---|
model | string | จำเป็น model id ของโมเดล open-weight ที่โฮสต์ |
messages | array | จำเป็น ข้อความในรูปแบบ Anthropic Messages นับบล็อก text, image, tool_use และ tool_result |
system | string | array | system prompt: สตริงหรืออาร์เรย์ของ text block |
tools | array | นิยาม tool ที่มี name, description และ input_schema |
รับไว้เพื่อความเข้ากันได้ โดยไม่มีผลต่อการนับ: tool_choice, max_tokens, temperature, top_p, stop_sequences, stream, thinking คุณส่ง body ของคำขอจริงได้โดยไม่ต้องแก้
คำตอบคือออบเจ็กต์ JSON ที่มีฟิลด์เหล่านี้:
| ฟิลด์ | ประเภท | คำอธิบาย |
|---|---|---|
input_tokens | integer | tokens ของอินพุตทั้งหมด: system prompt, messages, tools และรูปภาพ |
โมเดลที่รองรับ
ทั้งสอง endpoint นับให้โมเดล open-weight ที่โฮสต์ GET /v1/models ระบุ /v1/tokenize และ /v1/messages/count_tokens ใน endpoints ของแต่ละโมเดลที่รองรับ ค่า model อื่นใด รวมทั้ง id ของ Shannon จะได้รับ 400
DeepSeek-V4-Pro-0813-3BIT-REAPGLM-5.2-3BIT-REAPKimi-K3-3BIT-REAPNemotron3Ultra-3BIT-REAPMiniMax-M3-3BIT-REAPDeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAPKimi-K2.6-W4A16-AUTOROUND-REAPLaguna-S-2.1-W4A16-AUTOROUND-REAPinkling-W4A16-AUTOROUND-REAPMiMo-V2.5-Pro-W8A16MiMo-V2.5-W8A16Hy3-W8A16
สำหรับโมเดล Shannon ให้อ่านจำนวน tokens จากออบเจ็กต์ usage ของคำตอบ
การนับทำอย่างไร
แต่ละโมเดลถูกนับด้วย tokenizer และ chat template ของตัวเอง ไม่ใช้การประมาณจากจำนวนตัวอักษรหรือคำ
| สิ่งที่นับ | กฎ |
|---|---|
| ข้อความ | tokens ของสตริงตามที่ส่ง สตริงว่างนับเป็น 0 |
| ข้อความ | messages และ tools ถูกจัดวางด้วย chat template ของโมเดลเอง จนถึงจุดที่คำตอบเริ่ม และนับ prompt ทั้งหมดนั้น |
| บทบาท (role) | นับข้อความ system, user, assistant และ tool โดย developer นับเป็น system ข้อความที่ไม่มีเนื้อหาและไม่มีการเรียก tool จะไม่เพิ่มอะไร |
| การเรียก tool และผลลัพธ์ | การเรียก tool ของรอบ assistant ก่อนหน้าและผลลัพธ์ของมันเป็นส่วนหนึ่งของการนับ ในทั้งสอง endpoint |
| รูปภาพ | รูปภาพที่ส่งภายใน body (base64 หรือ data: URL) เพิ่มหนึ่ง token ต่อแพตช์ขนาด 28 × 28 พิกเซล: ceil(width / 28) × ceil(height / 28) รูปภาพที่ระบุเป็น URL แบบ http(s) จะไม่ถูกดาวน์โหลดโดย endpoint เหล่านี้ และนับเป็น 1,024 |
ตัวอย่าง: รูปภาพขนาด 1,024 × 768 พิกเซลนับเป็น ceil(1024 / 28) × ceil(768 / 28) = 37 × 28 = 1,036 tokens
การนับและสิ่งที่คำขอถูกคิดเงิน
การนับของทั้งคำขอทำแบบเดียวกับการนับอินพุตของคำขอจริงที่มีโมเดล messages และ tools เดียวกัน คำตอบรายงานตัวเลขนั้นเป็น usage.prompt_tokens บน Chat Completions เป็น usage.input_tokens บน Responses และเป็น usage.input_tokens บวก usage.cache_read_input_tokens บน Messages
- จำนวนที่นับคืออินพุตก่อนส่วนลดของอินพุตที่แคชไว้ คำขอจริงอาจอ่านอินพุตบางส่วนจากแคชและคิดเงินส่วนนั้นตามอัตราแคช การแคชพรอมต์
- รูปภาพที่ระบุเป็น URL แบบ
http(s)นับเป็น 1,024 ที่นี่ คำขอจริงดาวน์โหลดรูปภาพและนับจากขนาดพิกเซล ตัวเลขสองค่าจึงอาจต่างกัน ส่งรูปภาพเป็น base64 เพื่อให้ได้ตัวเลขเดียวกัน - เอาต์พุตไม่รวมอยู่ในการนับ คำตอบของคำขอจริงถูกคิดเงินเป็น output tokens เพิ่มเติม รวมการให้เหตุผลด้วย
- การนับ
textไม่มีการจัดรูปแบบแชต ใช้วัดเอกสารหรือส่วนของ prompt และใช้รูปแบบmessagesเพื่อวัดคำขอ
หากต้องการแปลงจำนวนที่นับเป็นค่าใช้จ่าย ให้คูณด้วยราคาอินพุตต่อ 1M tokens ของโมเดล โมเดลและราคา
ขีดจำกัด
| ขีดจำกัด | ค่า | เมื่อเกิน |
|---|---|---|
ความยาวของ text | 4,000,000 ไบต์ (UTF-8) | 413 พร้อมข้อความ text too long |
| เนื้อหาคำขอ (request body) | 32 MiB | 413 |
| ต่อคำขอ | หนึ่งข้อความหรือหนึ่งบทสนทนา | หากต้องการนับหลายข้อความ ให้ส่งหนึ่งคำขอต่อหนึ่งข้อความ |
การเรียกเพื่อนับไม่ถูกนับรวมในขีดจำกัด 120 คำขอต่อนาที ขีดจำกัดและยอดคงเหลือ
ข้อผิดพลาด
| สถานะ | ประเภท | ข้อความ | เมื่อใด |
|---|---|---|---|
400 | invalid_request_error | tokenize is available for the hosted open models; unknown model: <model> | /v1/tokenize ที่ model ไม่ใช่ id ของโมเดล open-weight ที่โฮสต์ |
400 | invalid_request_error | count_tokens is available for the hosted open models; unknown model: <model> | /v1/messages/count_tokens ที่ model ไม่ใช่ id ของโมเดล open-weight ที่โฮสต์ หรือไม่มี model |
400 | invalid_request_error | send `text` or `messages` | /v1/tokenize ที่ไม่มีทั้ง text และ messages |
401 | authentication_error | Missing authentication / Invalid API key | ไม่ได้ส่งคีย์ หรือคีย์ไม่ถูกต้อง |
413 | invalid_request_error | text too long | text ยาวกว่า 4,000,000 ไบต์ body ที่ใหญ่กว่า 32 MiB ก็ได้รับ 413 เช่นกัน |
415 | invalid_request_error | Expected request with `Content-Type: application/json` | คำขอไม่มี content type เป็น JSON |
422 | invalid_request_error | Failed to deserialize the JSON body into the target type: … | ขาดฟิลด์ที่จำเป็น (model บน /v1/tokenize, messages บน /v1/messages/count_tokens) หรือฟิลด์มีชนิดไม่ถูกต้อง |
503 | api_error | token counting is temporarily unavailable for this model | ขณะนี้นับสำหรับโมเดลนี้ไม่ได้ ให้ลองใหม่ภายหลัง |
/v1/tokenize ส่งคืนข้อผิดพลาดในรูปแบบ OpenAI บน /v1/messages/count_tokens ข้อผิดพลาดของ endpoint เอง (400 สำหรับโมเดล, 503) มาในรูปแบบ Anthropic ส่วน 401, 413, 415 และ 422 มาในรูปแบบ OpenAI ให้อ่านรหัสสถานะก่อน แล้วจึง error.type และ error.message ซึ่งมีในทั้งสองรูปแบบ
{
"error": {
"type": "invalid_request_error",
"message": "tokenize is available for the hosted open models; unknown model: shannon-3"
}
} {
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "count_tokens is available for the hosted open models; unknown model: shannon-3"
}
}