အကြောင်းအရာသို့ ကျော်သွားရန်
Token ရေတွက်ခြင်း

Token ရေတွက်ခြင်း

စာသားတစ်ခု သို့မဟုတ် request တစ်ခုလုံး၏ tokens ကို မပို့မီ ရေတွက်ပါ။

POST https://api.shannon-ai.com/v1/tokenize

POST https://api.shannon-ai.com/v1/messages/count_tokens

endpoint နှစ်ခုလုံးသည် သင်အမည်ပေးသော model ၏ tokenizer ဖြင့် ရေတွက်ပြီး model မ run ပါ။ ၎င်းတို့သည် host လုပ်ထားသော open-weight models များကို လွှမ်းခြုံသည်။ /v1/tokenize သည် သာမန်စာသား သို့မဟုတ် Chat Completions စကားပြောဆိုမှုကို လက်ခံသည်။ /v1/messages/count_tokens သည် Anthropic Messages format ဖြင့် request ကို လက်ခံပြီး Anthropic SDK နှင့် Claude Code ခေါ်သော call ဖြစ်သည်။

ရေတွက်ခြင်းသည် အခမဲ့ဖြစ်သည်။ call တစ်ခုတွင် သင့် API key လိုအပ်ပြီး သင့်လက်ကျန်ငွေမှ ဘာမျှ မနုတ်ဘဲ သင့် usage log တွင်လည်း မပေါ်ပါ။

စာသားတစ်ခုကို ရေတွက်ခြင်း

model နှင့် text ကို ပို့ပါ။ စာသားကို chat formatting မပါဘဲ ရှိသည့်အတိုင်း ရေတွက်သည်။

import requests

response = requests.post(
    "https://api.shannon-ai.com/v1/tokenize",
    headers={"Authorization": "Bearer YOUR_API_KEY"},
    json={
        "model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
        "text": "Hello, world",
    },
)
print(response.json()["tokens"])
200 ပြန်ကြားချက်
{
  "model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
  "tokens": 3
}

ဤစာမျက်နှာရှိ ပြန်ကြားချက်များရှိ နံပါတ်များသည် ဥပမာများ ဖြစ်သည်။ တူညီသောစာသားသည် model ကွဲလျှင် ရေတွက်မှု ကွဲသည်။

Chat request တစ်ခုကို ရေတွက်ခြင်း

/v1/chat/completions သို့ ပို့မည့်အတိုင်း model နှင့် messages ကို ပို့ပါ၊ request တွင် ရှိလျှင် tools ကိုပါ ပို့ပါ။ ပြန်ကြားချက်သည် input တစ်ခုလုံး၏ အရွယ်အစား ဖြစ်သည်။

import requests

request = {
    "model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
    "messages": [
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "What is the weather in Paris?"},
    ],
    "tools": [
        {
            "type": "function",
            "function": {
                "name": "get_weather",
                "description": "Current weather for a city",
                "parameters": {
                    "type": "object",
                    "properties": {"city": {"type": "string"}},
                    "required": ["city"],
                },
            },
        }
    ],
}

response = requests.post(
    "https://api.shannon-ai.com/v1/tokenize",
    headers={"Authorization": "Bearer YOUR_API_KEY"},
    json=request,
)
print(response.json()["tokens"])
200 ပြန်ကြားချက်
{
  "model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
  "tokens": 164
}

/v1/tokenize ၏ အချက်အလက်များ

Field Type ဖော်ပြချက်
model string လိုအပ်သည်။ host လုပ်ထားသော open-weight model id တစ်ခု။ စာလုံးအကြီးအသေးကို တူညီစွာ သဘောထားသည်။
text string chat formatting မပါဘဲ ရှိသည့်အတိုင်း ရေတွက်မည့် စာသား။ 4,000,000 bytes အထိ။ text သို့မဟုတ် messages ကို ပို့ပါ၊ နှစ်ခုလုံးပါလျှင် text ကို ရေတွက်သည်။
messages array Chat Completions format ဖြင့် chat messages။ ၎င်းတို့ကို request တစ်ခု၏ input အပြည့်အဝအဖြစ် ရေတွက်သည်- message တိုင်းကို model ၏ chat template က ၎င်းပတ်ပတ်လည်တွင် ထည့်သော formatting နှင့်အတူ။
tools array ရေတွက်မှုတွင် ထည့်သွင်းမည့် tool အဓိပ္ပာယ်ဖွင့်ဆိုချက်များ။ messages နှင့်အတူ သုံးသည်။

ပြန်ကြားချက်သည် ဤ field များပါသော JSON object တစ်ခု ဖြစ်သည်-

Field Type ဖော်ပြချက်
model string ရေတွက်မှုကို ပြုလုပ်ခဲ့သော model id၊ ၎င်း၏ ထုတ်ပြန်ထားသော စာလုံးပေါင်းဖြင့်။
tokens integer text ဖြင့်- စာသား၏ tokens။ messages ဖြင့်- ပုံများအပါအဝင် input တစ်ခုလုံး၏ tokens။

Messages request တစ်ခုကို ရေတွက်ခြင်း

/v1/messages သို့ ပို့မည့် body ကို ပို့ပါ- model၊ messages နှင့် သုံးလျှင် system နှင့် tools။ ရုံးသုံး Anthropic SDK များသည် ဤ endpoint ကို messages.count_tokens မှတစ်ဆင့် ခေါ်သည်။

import anthropic

client = anthropic.Anthropic(
    api_key="YOUR_API_KEY",
    base_url="https://api.shannon-ai.com",
)

count = client.messages.count_tokens(
    model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
    system="You are a concise assistant.",
    messages=[
        {"role": "user", "content": "Summarise the attached report."}
    ],
)
print(count.input_tokens)
200 ပြန်ကြားချက်
{
  "input_tokens": 21
}

/v1/messages/count_tokens သို့ ပေးပို့ရမည့် အချက်အလက်များ

Field Type ဖော်ပြချက်
model string လိုအပ်သည်။ host လုပ်ထားသော open-weight model ၏ id တစ်ခု ဖြစ်သည်။
messages array လိုအပ်သည်။ Anthropic Messages format ဖြင့် ရေးထားသော messages များ။ text၊ image၊ tool_use နှင့် tool_result block များကို ရေတွက်သည်။
system string | array system prompt ဖြစ်ပြီး string တစ်ခု သို့မဟုတ် text block များ၏ array ဖြစ်သည်။
tools array name၊ description နှင့် input_schema ပါသော tool အဓိပ္ပာယ်ဖွင့်ဆိုချက်များ။

ကိုက်ညီမှုအတွက် လက်ခံသော်လည်း ရေတွက်မှုကို မသက်ရောက်ပါ- tool_choice, max_tokens, temperature, top_p, stop_sequences, stream, thinking။ အမှန်တကယ် request ၏ body ကို မပြောင်းလဲဘဲ ပေးနိုင်သည်။

ပြန်ကြားချက်သည် ဤ field များပါသော JSON object တစ်ခု ဖြစ်သည်-

Field Type ဖော်ပြချက်
input_tokens integer system prompt၊ messages၊ tools နှင့် ပုံများ အပါအဝင် input တစ်ခုလုံး၏ token အရေအတွက်။

ပံ့ပိုးထားသော models

endpoint နှစ်ခုလုံးသည် host လုပ်ထားသော open-weight models များအတွက် ရေတွက်ပေးသည်။ GET /v1/models သည် ၎င်းတို့ကို ပံ့ပိုးသော model တစ်ခုစီ၏ endpoints တွင် /v1/tokenize နှင့် /v1/messages/count_tokens ကို စာရင်းပြသည်။ Shannon id များအပါအဝင် အခြား model တန်ဖိုးတိုင်းကို 400 ဖြင့် ဖြေသည်။

  • DeepSeek-V4-Pro-0813-3BIT-REAP
  • GLM-5.2-3BIT-REAP
  • Kimi-K3-3BIT-REAP
  • Nemotron3Ultra-3BIT-REAP
  • MiniMax-M3-3BIT-REAP
  • DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP
  • Kimi-K2.6-W4A16-AUTOROUND-REAP
  • Laguna-S-2.1-W4A16-AUTOROUND-REAP
  • inkling-W4A16-AUTOROUND-REAP
  • MiMo-V2.5-Pro-W8A16
  • MiMo-V2.5-W8A16
  • Hy3-W8A16

Shannon model အတွက် token ရေတွက်မှုများကို ပြန်ကြားချက်၏ usage object မှ ဖတ်ပါ။

ရေတွက်မှုကို မည်သို့ ပြုလုပ်သည်

Model တစ်ခုစီကို ၎င်း၏ ကိုယ်ပိုင် tokenizer နှင့် ကိုယ်ပိုင် chat template ဖြင့် ရေတွက်သည်။ စာလုံး သို့မဟုတ် စကားလုံးအရေအတွက်မှ ခန့်မှန်းချက်ကို မသုံးပါ။

ရေတွက်သည့်အရာ စည်းမျဉ်း
စာသားတစ်ခု ပို့သည့်အတိုင်း string ၏ tokens။ string အလွတ်ကို 0 အဖြစ် ရေတွက်သည်။
Messages messages နှင့် tools များကို ပြန်ကြားချက် မစတင်မီအထိ model ၏ ကိုယ်ပိုင် chat template ဖြင့် စီစဉ်ပြီး prompt တစ်ခုလုံးကို ရေတွက်သည်။
Roles system၊ user၊ assistant နှင့် tool messages များကို ရေတွက်သည်။ developer ကို system အဖြစ် ရေတွက်သည်။ content မပါ tool call လည်း မပါသော message သည် ဘာမျှ မထည့်ပါ။
Tool calls နှင့် ရလဒ်များ ယခင် assistant turn များ၏ tool calls နှင့် ၎င်းတို့၏ ရလဒ်များသည် endpoint နှစ်ခုလုံးတွင် ရေတွက်မှု၏ အစိတ်အပိုင်းဖြစ်သည်။
ပုံများ body အတွင်း (base64 သို့မဟုတ် data: URL) ပို့သော ပုံသည် 28 × 28 pixel patch တစ်ခုလျှင် token တစ်ခု ထပ်ထည့်သည်- ceil(width / 28) × ceil(height / 28)။ http(s) URL အဖြစ် ပေးထားသော ပုံကို ဤ endpoint များက download မလုပ်ဘဲ 1,024 အဖြစ် ရေတွက်သည်။

ဥပမာ- 1,024 × 768 pixel ပုံတစ်ပုံကို ceil(1024 / 28) × ceil(768 / 28) = 37 × 28 = 1,036 tokens အဖြစ် ရေတွက်သည်။

ရေတွက်မှုနှင့် request တစ်ခုအတွက် ကျသင့်ငွေ

request တစ်ခုလုံး၏ ရေတွက်မှုကို model၊ messages နှင့် tools တူညီသော အမှန်တကယ် request ၏ input ရေတွက်မှုအတိုင်း တူညီစွာ ပြုလုပ်သည်။ ပြန်ကြားချက်က ထိုနံပါတ်ကို Chat Completions တွင် usage.prompt_tokens အဖြစ်၊ Responses တွင် usage.input_tokens အဖြစ်၊ Messages တွင် usage.input_tokens ပေါင်း usage.cache_read_input_tokens အဖြစ် ဖော်ပြသည်။

  • ရေတွက်မှုသည် cached-input လျှော့ဈေးမတိုင်မီ input ဖြစ်သည်။ အမှန်တကယ် request သည် ထို input အချို့ကို cache မှ ဖတ်ပြီး ထိုအပိုင်းကို cached နှုန်းဖြင့် ကျသင့်ငွေတောင်းနိုင်သည်။ Prompt caching
  • http(s) URL အဖြစ် ပေးထားသော ပုံကို ဤနေရာတွင် 1,024 အဖြစ် ရေတွက်သည်။ အမှန်တကယ် request သည် ပုံကို download လုပ်ပြီး pixel အရွယ်အစားမှ ရေတွက်သဖြင့် နံပါတ်နှစ်ခု ကွဲနိုင်သည်။ တူညီသောနံပါတ်ရရန် ပုံကို base64 အဖြစ် ပို့ပါ။
  • Output သည် ရေတွက်မှု၏ အစိတ်အပိုင်း မဟုတ်ပါ။ အမှန်တကယ် request ၏ ပြန်ကြားချက်ကို reasoning အပါအဝင် output tokens အဖြစ် ထပ်ပေါင်း၍ ကျသင့်ငွေတောင်းသည်။
  • text ရေတွက်မှုတွင် chat formatting မပါပါ။ စာရွက်စာတမ်း သို့မဟုတ် prompt အပိုင်းတစ်ခုကို တိုင်းရန် ၎င်းကို သုံးပြီး request တစ်ခုကို တိုင်းရန် messages ပုံစံကို သုံးပါ။

ရေတွက်မှုကို ကုန်ကျစရိတ်သို့ ပြောင်းရန် model ၏ tokens 1M လျှင် input ဈေးနှုန်းဖြင့် မြှောက်ပါ။ Model များနှင့် ဈေးနှုန်း

ကန့်သတ်ချက်များ

ကန့်သတ်ချက် တန်ဖိုး ကျော်လွန်လျှင်
text ၏ အရှည် 4,000,000 bytes (UTF-8) 413 နှင့် မက်ဆေ့ချ် text too long
Request body 32 MiB 413
Request တစ်ခုလျှင် စာသားတစ်ခု သို့မဟုတ် စကားပြောဆိုမှုတစ်ခု စာသားများစွာကို ရေတွက်ရန် စာသားတစ်ခုလျှင် request တစ်ခု ပို့ပါ။

ရေတွက်သော calls များကို တစ်မိနစ်လျှင် request 120 ကန့်သတ်ချက်တွင် မရေတွက်ပါ။ ကန့်သတ်ချက်များနှင့် လက်ကျန်ငွေ

အမှားများ

Status Type မက်ဆေ့ချ် ဖြစ်ပေါ်ချိန်
400 invalid_request_error tokenize is available for the hosted open models; unknown model: <model> host လုပ်ထားသော open-weight id မဟုတ်သည့် model ပါသော /v1/tokenize။
400 invalid_request_error count_tokens is available for the hosted open models; unknown model: <model> host လုပ်ထားသော open-weight id မဟုတ်သည့် model ပါသော သို့မဟုတ် model မပါသော /v1/messages/count_tokens။
400 invalid_request_error send `text` or `messages` text မပါ၊ messages လည်း မပါသော /v1/tokenize။
401 authentication_error Missing authentication / Invalid API key key မပို့ခဲ့ပါ၊ သို့မဟုတ် key မမှန်ကန်ပါ။
413 invalid_request_error text too long text သည် 4,000,000 bytes ထက် ပိုရှည်သည်။ 32 MiB ထက်ပိုသော body ကိုလည်း 413 ဖြင့် ဖြေသည်။
415 invalid_request_error Expected request with `Content-Type: application/json` request တွင် JSON ဖြစ်ကြောင်း ဖော်ပြသော content type မပါပါ။
422 invalid_request_error Failed to deserialize the JSON body into the target type: … လိုအပ်သော field မပါ (/v1/tokenize တွင် model၊ /v1/messages/count_tokens တွင် messages) သို့မဟုတ် field တစ်ခု၏ type မှားနေသည်။
503 api_error token counting is temporarily unavailable for this model ယခုအချိန်တွင် ဤ model အတွက် ရေတွက်၍မရပါ။ နောက်မှ ထပ်ကြိုးစားပါ။

/v1/tokenize သည် အမှားများကို OpenAI ပုံစံဖြင့် ပြန်ပေးသည်။ /v1/messages/count_tokens တွင် endpoint ကိုယ်တိုင်၏ အမှားများ (model အတွက် 400၊ 503) သည် Anthropic ပုံစံဖြင့် လာပြီး 401၊ 413၊ 415 နှင့် 422 တို့သည် OpenAI ပုံစံဖြင့် လာသည်။ status code ကို အရင်ဖတ်ပြီး ပုံစံနှစ်ခုလုံးတွင် ပါသော error.type နှင့် error.message ကို ဖတ်ပါ။

400 /v1/tokenize
{
  "error": {
    "type": "invalid_request_error",
    "message": "tokenize is available for the hosted open models; unknown model: shannon-3"
  }
}
400 /v1/messages/count_tokens
{
  "type": "error",
  "error": {
    "type": "invalid_request_error",
    "message": "count_tokens is available for the hosted open models; unknown model: shannon-3"
  }
}