ข้ามไปยังเนื้อหา
การนับ token

การนับ token

นับ tokens ของข้อความหรือของทั้งคำขอก่อนที่คุณจะส่ง

POST https://api.shannon-ai.com/v1/tokenize

POST https://api.shannon-ai.com/v1/messages/count_tokens

ทั้งสอง endpoint นับด้วย tokenizer ของโมเดลที่คุณระบุ และไม่มีโมเดลทำงาน ครอบคลุมโมเดล open-weight ที่โฮสต์ /v1/tokenize รับข้อความธรรมดาหรือบทสนทนา Chat Completions /v1/messages/count_tokens รับคำขอในรูปแบบ Anthropic Messages ซึ่งเป็นการเรียกที่ Anthropic SDK และ Claude Code ทำ

การนับไม่มีค่าใช้จ่าย การเรียกต้องใช้ API key ของคุณ ไม่หักจากยอดคงเหลือ และไม่ปรากฏในบันทึกการใช้งานของคุณ

นับข้อความ

ส่ง model และ text ข้อความจะถูกนับตามที่เป็น โดยไม่มีการจัดรูปแบบแชตล้อมรอบ

import requests

response = requests.post(
    "https://api.shannon-ai.com/v1/tokenize",
    headers={"Authorization": "Bearer YOUR_API_KEY"},
    json={
        "model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
        "text": "Hello, world",
    },
)
print(response.json()["tokens"])
200 คำตอบ
{
  "model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
  "tokens": 3
}

ตัวเลขในคำตอบของหน้านี้เป็นตัวอย่าง ข้อความเดียวกันให้จำนวนที่นับต่างกันบนโมเดลที่ต่างกัน

นับคำขอแชต

ส่ง model และ messages พร้อม tools เมื่อคำขอมี เหมือนที่จะส่งไปยัง /v1/chat/completions ทุกประการ คำตอบคือขนาดของอินพุตทั้งหมด

import requests

request = {
    "model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
    "messages": [
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "What is the weather in Paris?"},
    ],
    "tools": [
        {
            "type": "function",
            "function": {
                "name": "get_weather",
                "description": "Current weather for a city",
                "parameters": {
                    "type": "object",
                    "properties": {"city": {"type": "string"}},
                    "required": ["city"],
                },
            },
        }
    ],
}

response = requests.post(
    "https://api.shannon-ai.com/v1/tokenize",
    headers={"Authorization": "Bearer YOUR_API_KEY"},
    json=request,
)
print(response.json()["tokens"])
200 คำตอบ
{
  "model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
  "tokens": 164
}

ฟิลด์ของ /v1/tokenize

ฟิลด์ ประเภท คำอธิบาย
model string จำเป็น model id ของโมเดล open-weight ที่โฮสต์ ตัวพิมพ์ใหญ่และเล็กถือเป็นอย่างเดียวกัน
text string ข้อความที่นับตามที่เป็น โดยไม่มีการจัดรูปแบบแชต ไม่เกิน 4,000,000 ไบต์ ส่ง text หรือ messages เมื่อมีทั้งสอง จะนับ text
messages array ข้อความแชตในรูปแบบ Chat Completions ถูกนับเป็นอินพุตเต็มของคำขอ: ทุกข้อความพร้อมการจัดรูปแบบที่ chat template ของโมเดลใส่รอบข้อความนั้น
tools array นิยาม tool ที่ต้องรวมในการนับ ใช้ร่วมกับ messages

คำตอบคือออบเจ็กต์ JSON ที่มีฟิลด์เหล่านี้:

ฟิลด์ ประเภท คำอธิบาย
model string model id ที่ใช้นับ ในรูปที่เผยแพร่
tokens integer เมื่อใช้ text: tokens ของข้อความ เมื่อใช้ messages: tokens ของอินพุตทั้งหมด รวมรูปภาพ

นับคำขอ Messages

ส่ง body เดียวกับที่ส่งไป /v1/messages: model, messages และ system กับ tools เมื่อใช้ Anthropic SDK อย่างเป็นทางการเรียก endpoint นี้ผ่าน messages.count_tokens

import anthropic

client = anthropic.Anthropic(
    api_key="YOUR_API_KEY",
    base_url="https://api.shannon-ai.com",
)

count = client.messages.count_tokens(
    model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
    system="You are a concise assistant.",
    messages=[
        {"role": "user", "content": "Summarise the attached report."}
    ],
)
print(count.input_tokens)
200 คำตอบ
{
  "input_tokens": 21
}

ฟิลด์ที่ใช้กับ endpoint สำหรับนับ token แบบ Messages (/v1/messages/count_tokens)

ฟิลด์ ประเภท คำอธิบาย
model string จำเป็น model id ของโมเดล open-weight ที่โฮสต์
messages array จำเป็น ข้อความในรูปแบบ Anthropic Messages นับบล็อก text, image, tool_use และ tool_result
system string | array system prompt: สตริงหรืออาร์เรย์ของ text block
tools array นิยาม tool ที่มี name, description และ input_schema

รับไว้เพื่อความเข้ากันได้ โดยไม่มีผลต่อการนับ: tool_choice, max_tokens, temperature, top_p, stop_sequences, stream, thinking คุณส่ง body ของคำขอจริงได้โดยไม่ต้องแก้

คำตอบคือออบเจ็กต์ JSON ที่มีฟิลด์เหล่านี้:

ฟิลด์ ประเภท คำอธิบาย
input_tokens integer tokens ของอินพุตทั้งหมด: system prompt, messages, tools และรูปภาพ

โมเดลที่รองรับ

ทั้งสอง endpoint นับให้โมเดล open-weight ที่โฮสต์ GET /v1/models ระบุ /v1/tokenize และ /v1/messages/count_tokens ใน endpoints ของแต่ละโมเดลที่รองรับ ค่า model อื่นใด รวมทั้ง id ของ Shannon จะได้รับ 400

  • DeepSeek-V4-Pro-0813-3BIT-REAP
  • GLM-5.2-3BIT-REAP
  • Kimi-K3-3BIT-REAP
  • Nemotron3Ultra-3BIT-REAP
  • MiniMax-M3-3BIT-REAP
  • DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP
  • Kimi-K2.6-W4A16-AUTOROUND-REAP
  • Laguna-S-2.1-W4A16-AUTOROUND-REAP
  • inkling-W4A16-AUTOROUND-REAP
  • MiMo-V2.5-Pro-W8A16
  • MiMo-V2.5-W8A16
  • Hy3-W8A16

สำหรับโมเดล Shannon ให้อ่านจำนวน tokens จากออบเจ็กต์ usage ของคำตอบ

การนับทำอย่างไร

แต่ละโมเดลถูกนับด้วย tokenizer และ chat template ของตัวเอง ไม่ใช้การประมาณจากจำนวนตัวอักษรหรือคำ

สิ่งที่นับ กฎ
ข้อความ tokens ของสตริงตามที่ส่ง สตริงว่างนับเป็น 0
ข้อความ messages และ tools ถูกจัดวางด้วย chat template ของโมเดลเอง จนถึงจุดที่คำตอบเริ่ม และนับ prompt ทั้งหมดนั้น
บทบาท (role) นับข้อความ system, user, assistant และ tool โดย developer นับเป็น system ข้อความที่ไม่มีเนื้อหาและไม่มีการเรียก tool จะไม่เพิ่มอะไร
การเรียก tool และผลลัพธ์ การเรียก tool ของรอบ assistant ก่อนหน้าและผลลัพธ์ของมันเป็นส่วนหนึ่งของการนับ ในทั้งสอง endpoint
รูปภาพ รูปภาพที่ส่งภายใน body (base64 หรือ data: URL) เพิ่มหนึ่ง token ต่อแพตช์ขนาด 28 × 28 พิกเซล: ceil(width / 28) × ceil(height / 28) รูปภาพที่ระบุเป็น URL แบบ http(s) จะไม่ถูกดาวน์โหลดโดย endpoint เหล่านี้ และนับเป็น 1,024

ตัวอย่าง: รูปภาพขนาด 1,024 × 768 พิกเซลนับเป็น ceil(1024 / 28) × ceil(768 / 28) = 37 × 28 = 1,036 tokens

การนับและสิ่งที่คำขอถูกคิดเงิน

การนับของทั้งคำขอทำแบบเดียวกับการนับอินพุตของคำขอจริงที่มีโมเดล messages และ tools เดียวกัน คำตอบรายงานตัวเลขนั้นเป็น usage.prompt_tokens บน Chat Completions เป็น usage.input_tokens บน Responses และเป็น usage.input_tokens บวก usage.cache_read_input_tokens บน Messages

  • จำนวนที่นับคืออินพุตก่อนส่วนลดของอินพุตที่แคชไว้ คำขอจริงอาจอ่านอินพุตบางส่วนจากแคชและคิดเงินส่วนนั้นตามอัตราแคช การแคชพรอมต์
  • รูปภาพที่ระบุเป็น URL แบบ http(s) นับเป็น 1,024 ที่นี่ คำขอจริงดาวน์โหลดรูปภาพและนับจากขนาดพิกเซล ตัวเลขสองค่าจึงอาจต่างกัน ส่งรูปภาพเป็น base64 เพื่อให้ได้ตัวเลขเดียวกัน
  • เอาต์พุตไม่รวมอยู่ในการนับ คำตอบของคำขอจริงถูกคิดเงินเป็น output tokens เพิ่มเติม รวมการให้เหตุผลด้วย
  • การนับ text ไม่มีการจัดรูปแบบแชต ใช้วัดเอกสารหรือส่วนของ prompt และใช้รูปแบบ messages เพื่อวัดคำขอ

หากต้องการแปลงจำนวนที่นับเป็นค่าใช้จ่าย ให้คูณด้วยราคาอินพุตต่อ 1M tokens ของโมเดล โมเดลและราคา

ขีดจำกัด

ขีดจำกัด ค่า เมื่อเกิน
ความยาวของ text 4,000,000 ไบต์ (UTF-8) 413 พร้อมข้อความ text too long
เนื้อหาคำขอ (request body) 32 MiB 413
ต่อคำขอ หนึ่งข้อความหรือหนึ่งบทสนทนา หากต้องการนับหลายข้อความ ให้ส่งหนึ่งคำขอต่อหนึ่งข้อความ

การเรียกเพื่อนับไม่ถูกนับรวมในขีดจำกัด 120 คำขอต่อนาที ขีดจำกัดและยอดคงเหลือ

ข้อผิดพลาด

สถานะ ประเภท ข้อความ เมื่อใด
400 invalid_request_error tokenize is available for the hosted open models; unknown model: <model> /v1/tokenize ที่ model ไม่ใช่ id ของโมเดล open-weight ที่โฮสต์
400 invalid_request_error count_tokens is available for the hosted open models; unknown model: <model> /v1/messages/count_tokens ที่ model ไม่ใช่ id ของโมเดล open-weight ที่โฮสต์ หรือไม่มี model
400 invalid_request_error send `text` or `messages` /v1/tokenize ที่ไม่มีทั้ง text และ messages
401 authentication_error Missing authentication / Invalid API key ไม่ได้ส่งคีย์ หรือคีย์ไม่ถูกต้อง
413 invalid_request_error text too long text ยาวกว่า 4,000,000 ไบต์ body ที่ใหญ่กว่า 32 MiB ก็ได้รับ 413 เช่นกัน
415 invalid_request_error Expected request with `Content-Type: application/json` คำขอไม่มี content type เป็น JSON
422 invalid_request_error Failed to deserialize the JSON body into the target type: … ขาดฟิลด์ที่จำเป็น (model บน /v1/tokenize, messages บน /v1/messages/count_tokens) หรือฟิลด์มีชนิดไม่ถูกต้อง
503 api_error token counting is temporarily unavailable for this model ขณะนี้นับสำหรับโมเดลนี้ไม่ได้ ให้ลองใหม่ภายหลัง

/v1/tokenize ส่งคืนข้อผิดพลาดในรูปแบบ OpenAI บน /v1/messages/count_tokens ข้อผิดพลาดของ endpoint เอง (400 สำหรับโมเดล, 503) มาในรูปแบบ Anthropic ส่วน 401, 413, 415 และ 422 มาในรูปแบบ OpenAI ให้อ่านรหัสสถานะก่อน แล้วจึง error.type และ error.message ซึ่งมีในทั้งสองรูปแบบ

400 /v1/tokenize
{
  "error": {
    "type": "invalid_request_error",
    "message": "tokenize is available for the hosted open models; unknown model: shannon-3"
  }
}
400 /v1/messages/count_tokens
{
  "type": "error",
  "error": {
    "type": "invalid_request_error",
    "message": "count_tokens is available for the hosted open models; unknown model: shannon-3"
  }
}