Token ရေတွက်ခြင်း
စာသားတစ်ခု သို့မဟုတ် request တစ်ခုလုံး၏ tokens ကို မပို့မီ ရေတွက်ပါ။
POST https://api.shannon-ai.com/v1/tokenize
POST https://api.shannon-ai.com/v1/messages/count_tokens
endpoint နှစ်ခုလုံးသည် သင်အမည်ပေးသော model ၏ tokenizer ဖြင့် ရေတွက်ပြီး model မ run ပါ။ ၎င်းတို့သည် host လုပ်ထားသော open-weight models များကို လွှမ်းခြုံသည်။ /v1/tokenize သည် သာမန်စာသား သို့မဟုတ် Chat Completions စကားပြောဆိုမှုကို လက်ခံသည်။ /v1/messages/count_tokens သည် Anthropic Messages format ဖြင့် request ကို လက်ခံပြီး Anthropic SDK နှင့် Claude Code ခေါ်သော call ဖြစ်သည်။
ရေတွက်ခြင်းသည် အခမဲ့ဖြစ်သည်။ call တစ်ခုတွင် သင့် API key လိုအပ်ပြီး သင့်လက်ကျန်ငွေမှ ဘာမျှ မနုတ်ဘဲ သင့် usage log တွင်လည်း မပေါ်ပါ။
စာသားတစ်ခုကို ရေတွက်ခြင်း
model နှင့် text ကို ပို့ပါ။ စာသားကို chat formatting မပါဘဲ ရှိသည့်အတိုင်း ရေတွက်သည်။
import requests
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world",
},
)
print(response.json()["tokens"]) const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
text: "Hello, world",
}),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world"
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 3
} ဤစာမျက်နှာရှိ ပြန်ကြားချက်များရှိ နံပါတ်များသည် ဥပမာများ ဖြစ်သည်။ တူညီသောစာသားသည် model ကွဲလျှင် ရေတွက်မှု ကွဲသည်။
Chat request တစ်ခုကို ရေတွက်ခြင်း
/v1/chat/completions သို့ ပို့မည့်အတိုင်း model နှင့် messages ကို ပို့ပါ၊ request တွင် ရှိလျှင် tools ကိုပါ ပို့ပါ။ ပြန်ကြားချက်သည် input တစ်ခုလုံး၏ အရွယ်အစား ဖြစ်သည်။
import requests
request = {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"},
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
],
}
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json=request,
)
print(response.json()["tokens"]) const request = {
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
messages: [
{ role: "system", content: "You are a concise assistant." },
{ role: "user", content: "What is the weather in Paris?" },
],
tools: [
{
type: "function",
function: {
name: "get_weather",
description: "Current weather for a city",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
},
},
},
],
};
const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify(request),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}
]
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 164
} /v1/tokenize ၏ အချက်အလက်များ
| Field | Type | ဖော်ပြချက် |
|---|---|---|
model | string | လိုအပ်သည်။ host လုပ်ထားသော open-weight model id တစ်ခု။ စာလုံးအကြီးအသေးကို တူညီစွာ သဘောထားသည်။ |
text | string | chat formatting မပါဘဲ ရှိသည့်အတိုင်း ရေတွက်မည့် စာသား။ 4,000,000 bytes အထိ။ text သို့မဟုတ် messages ကို ပို့ပါ၊ နှစ်ခုလုံးပါလျှင် text ကို ရေတွက်သည်။ |
messages | array | Chat Completions format ဖြင့် chat messages။ ၎င်းတို့ကို request တစ်ခု၏ input အပြည့်အဝအဖြစ် ရေတွက်သည်- message တိုင်းကို model ၏ chat template က ၎င်းပတ်ပတ်လည်တွင် ထည့်သော formatting နှင့်အတူ။ |
tools | array | ရေတွက်မှုတွင် ထည့်သွင်းမည့် tool အဓိပ္ပာယ်ဖွင့်ဆိုချက်များ။ messages နှင့်အတူ သုံးသည်။ |
ပြန်ကြားချက်သည် ဤ field များပါသော JSON object တစ်ခု ဖြစ်သည်-
| Field | Type | ဖော်ပြချက် |
|---|---|---|
model | string | ရေတွက်မှုကို ပြုလုပ်ခဲ့သော model id၊ ၎င်း၏ ထုတ်ပြန်ထားသော စာလုံးပေါင်းဖြင့်။ |
tokens | integer | text ဖြင့်- စာသား၏ tokens။ messages ဖြင့်- ပုံများအပါအဝင် input တစ်ခုလုံး၏ tokens။ |
Messages request တစ်ခုကို ရေတွက်ခြင်း
/v1/messages သို့ ပို့မည့် body ကို ပို့ပါ- model၊ messages နှင့် သုံးလျှင် system နှင့် tools။ ရုံးသုံး Anthropic SDK များသည် ဤ endpoint ကို messages.count_tokens မှတစ်ဆင့် ခေါ်သည်။
import anthropic
client = anthropic.Anthropic(
api_key="YOUR_API_KEY",
base_url="https://api.shannon-ai.com",
)
count = client.messages.count_tokens(
model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system="You are a concise assistant.",
messages=[
{"role": "user", "content": "Summarise the attached report."}
],
)
print(count.input_tokens) import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
apiKey: "YOUR_API_KEY",
baseURL: "https://api.shannon-ai.com",
});
const count = await client.messages.countTokens({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system: "You are a concise assistant.",
messages: [
{ role: "user", content: "Summarise the attached report." },
],
});
console.log(count.input_tokens); curl https://api.shannon-ai.com/v1/messages/count_tokens \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"system": "You are a concise assistant.",
"messages": [
{"role": "user", "content": "Summarise the attached report."}
]
}' {
"input_tokens": 21
} /v1/messages/count_tokens သို့ ပေးပို့ရမည့် အချက်အလက်များ
| Field | Type | ဖော်ပြချက် |
|---|---|---|
model | string | လိုအပ်သည်။ host လုပ်ထားသော open-weight model ၏ id တစ်ခု ဖြစ်သည်။ |
messages | array | လိုအပ်သည်။ Anthropic Messages format ဖြင့် ရေးထားသော messages များ။ text၊ image၊ tool_use နှင့် tool_result block များကို ရေတွက်သည်။ |
system | string | array | system prompt ဖြစ်ပြီး string တစ်ခု သို့မဟုတ် text block များ၏ array ဖြစ်သည်။ |
tools | array | name၊ description နှင့် input_schema ပါသော tool အဓိပ္ပာယ်ဖွင့်ဆိုချက်များ။ |
ကိုက်ညီမှုအတွက် လက်ခံသော်လည်း ရေတွက်မှုကို မသက်ရောက်ပါ- tool_choice, max_tokens, temperature, top_p, stop_sequences, stream, thinking။ အမှန်တကယ် request ၏ body ကို မပြောင်းလဲဘဲ ပေးနိုင်သည်။
ပြန်ကြားချက်သည် ဤ field များပါသော JSON object တစ်ခု ဖြစ်သည်-
| Field | Type | ဖော်ပြချက် |
|---|---|---|
input_tokens | integer | system prompt၊ messages၊ tools နှင့် ပုံများ အပါအဝင် input တစ်ခုလုံး၏ token အရေအတွက်။ |
ပံ့ပိုးထားသော models
endpoint နှစ်ခုလုံးသည် host လုပ်ထားသော open-weight models များအတွက် ရေတွက်ပေးသည်။ GET /v1/models သည် ၎င်းတို့ကို ပံ့ပိုးသော model တစ်ခုစီ၏ endpoints တွင် /v1/tokenize နှင့် /v1/messages/count_tokens ကို စာရင်းပြသည်။ Shannon id များအပါအဝင် အခြား model တန်ဖိုးတိုင်းကို 400 ဖြင့် ဖြေသည်။
DeepSeek-V4-Pro-0813-3BIT-REAPGLM-5.2-3BIT-REAPKimi-K3-3BIT-REAPNemotron3Ultra-3BIT-REAPMiniMax-M3-3BIT-REAPDeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAPKimi-K2.6-W4A16-AUTOROUND-REAPLaguna-S-2.1-W4A16-AUTOROUND-REAPinkling-W4A16-AUTOROUND-REAPMiMo-V2.5-Pro-W8A16MiMo-V2.5-W8A16Hy3-W8A16
Shannon model အတွက် token ရေတွက်မှုများကို ပြန်ကြားချက်၏ usage object မှ ဖတ်ပါ။
ရေတွက်မှုကို မည်သို့ ပြုလုပ်သည်
Model တစ်ခုစီကို ၎င်း၏ ကိုယ်ပိုင် tokenizer နှင့် ကိုယ်ပိုင် chat template ဖြင့် ရေတွက်သည်။ စာလုံး သို့မဟုတ် စကားလုံးအရေအတွက်မှ ခန့်မှန်းချက်ကို မသုံးပါ။
| ရေတွက်သည့်အရာ | စည်းမျဉ်း |
|---|---|
| စာသားတစ်ခု | ပို့သည့်အတိုင်း string ၏ tokens။ string အလွတ်ကို 0 အဖြစ် ရေတွက်သည်။ |
| Messages | messages နှင့် tools များကို ပြန်ကြားချက် မစတင်မီအထိ model ၏ ကိုယ်ပိုင် chat template ဖြင့် စီစဉ်ပြီး prompt တစ်ခုလုံးကို ရေတွက်သည်။ |
| Roles | system၊ user၊ assistant နှင့် tool messages များကို ရေတွက်သည်။ developer ကို system အဖြစ် ရေတွက်သည်။ content မပါ tool call လည်း မပါသော message သည် ဘာမျှ မထည့်ပါ။ |
| Tool calls နှင့် ရလဒ်များ | ယခင် assistant turn များ၏ tool calls နှင့် ၎င်းတို့၏ ရလဒ်များသည် endpoint နှစ်ခုလုံးတွင် ရေတွက်မှု၏ အစိတ်အပိုင်းဖြစ်သည်။ |
| ပုံများ | body အတွင်း (base64 သို့မဟုတ် data: URL) ပို့သော ပုံသည် 28 × 28 pixel patch တစ်ခုလျှင် token တစ်ခု ထပ်ထည့်သည်- ceil(width / 28) × ceil(height / 28)။ http(s) URL အဖြစ် ပေးထားသော ပုံကို ဤ endpoint များက download မလုပ်ဘဲ 1,024 အဖြစ် ရေတွက်သည်။ |
ဥပမာ- 1,024 × 768 pixel ပုံတစ်ပုံကို ceil(1024 / 28) × ceil(768 / 28) = 37 × 28 = 1,036 tokens အဖြစ် ရေတွက်သည်။
ရေတွက်မှုနှင့် request တစ်ခုအတွက် ကျသင့်ငွေ
request တစ်ခုလုံး၏ ရေတွက်မှုကို model၊ messages နှင့် tools တူညီသော အမှန်တကယ် request ၏ input ရေတွက်မှုအတိုင်း တူညီစွာ ပြုလုပ်သည်။ ပြန်ကြားချက်က ထိုနံပါတ်ကို Chat Completions တွင် usage.prompt_tokens အဖြစ်၊ Responses တွင် usage.input_tokens အဖြစ်၊ Messages တွင် usage.input_tokens ပေါင်း usage.cache_read_input_tokens အဖြစ် ဖော်ပြသည်။
- ရေတွက်မှုသည် cached-input လျှော့ဈေးမတိုင်မီ input ဖြစ်သည်။ အမှန်တကယ် request သည် ထို input အချို့ကို cache မှ ဖတ်ပြီး ထိုအပိုင်းကို cached နှုန်းဖြင့် ကျသင့်ငွေတောင်းနိုင်သည်။ Prompt caching
http(s)URL အဖြစ် ပေးထားသော ပုံကို ဤနေရာတွင် 1,024 အဖြစ် ရေတွက်သည်။ အမှန်တကယ် request သည် ပုံကို download လုပ်ပြီး pixel အရွယ်အစားမှ ရေတွက်သဖြင့် နံပါတ်နှစ်ခု ကွဲနိုင်သည်။ တူညီသောနံပါတ်ရရန် ပုံကို base64 အဖြစ် ပို့ပါ။- Output သည် ရေတွက်မှု၏ အစိတ်အပိုင်း မဟုတ်ပါ။ အမှန်တကယ် request ၏ ပြန်ကြားချက်ကို reasoning အပါအဝင် output tokens အဖြစ် ထပ်ပေါင်း၍ ကျသင့်ငွေတောင်းသည်။
textရေတွက်မှုတွင် chat formatting မပါပါ။ စာရွက်စာတမ်း သို့မဟုတ် prompt အပိုင်းတစ်ခုကို တိုင်းရန် ၎င်းကို သုံးပြီး request တစ်ခုကို တိုင်းရန်messagesပုံစံကို သုံးပါ။
ရေတွက်မှုကို ကုန်ကျစရိတ်သို့ ပြောင်းရန် model ၏ tokens 1M လျှင် input ဈေးနှုန်းဖြင့် မြှောက်ပါ။ Model များနှင့် ဈေးနှုန်း
ကန့်သတ်ချက်များ
| ကန့်သတ်ချက် | တန်ဖိုး | ကျော်လွန်လျှင် |
|---|---|---|
text ၏ အရှည် | 4,000,000 bytes (UTF-8) | 413 နှင့် မက်ဆေ့ချ် text too long |
| Request body | 32 MiB | 413 |
| Request တစ်ခုလျှင် | စာသားတစ်ခု သို့မဟုတ် စကားပြောဆိုမှုတစ်ခု | စာသားများစွာကို ရေတွက်ရန် စာသားတစ်ခုလျှင် request တစ်ခု ပို့ပါ။ |
ရေတွက်သော calls များကို တစ်မိနစ်လျှင် request 120 ကန့်သတ်ချက်တွင် မရေတွက်ပါ။ ကန့်သတ်ချက်များနှင့် လက်ကျန်ငွေ
အမှားများ
| Status | Type | မက်ဆေ့ချ် | ဖြစ်ပေါ်ချိန် |
|---|---|---|---|
400 | invalid_request_error | tokenize is available for the hosted open models; unknown model: <model> | host လုပ်ထားသော open-weight id မဟုတ်သည့် model ပါသော /v1/tokenize။ |
400 | invalid_request_error | count_tokens is available for the hosted open models; unknown model: <model> | host လုပ်ထားသော open-weight id မဟုတ်သည့် model ပါသော သို့မဟုတ် model မပါသော /v1/messages/count_tokens။ |
400 | invalid_request_error | send `text` or `messages` | text မပါ၊ messages လည်း မပါသော /v1/tokenize။ |
401 | authentication_error | Missing authentication / Invalid API key | key မပို့ခဲ့ပါ၊ သို့မဟုတ် key မမှန်ကန်ပါ။ |
413 | invalid_request_error | text too long | text သည် 4,000,000 bytes ထက် ပိုရှည်သည်။ 32 MiB ထက်ပိုသော body ကိုလည်း 413 ဖြင့် ဖြေသည်။ |
415 | invalid_request_error | Expected request with `Content-Type: application/json` | request တွင် JSON ဖြစ်ကြောင်း ဖော်ပြသော content type မပါပါ။ |
422 | invalid_request_error | Failed to deserialize the JSON body into the target type: … | လိုအပ်သော field မပါ (/v1/tokenize တွင် model၊ /v1/messages/count_tokens တွင် messages) သို့မဟုတ် field တစ်ခု၏ type မှားနေသည်။ |
503 | api_error | token counting is temporarily unavailable for this model | ယခုအချိန်တွင် ဤ model အတွက် ရေတွက်၍မရပါ။ နောက်မှ ထပ်ကြိုးစားပါ။ |
/v1/tokenize သည် အမှားများကို OpenAI ပုံစံဖြင့် ပြန်ပေးသည်။ /v1/messages/count_tokens တွင် endpoint ကိုယ်တိုင်၏ အမှားများ (model အတွက် 400၊ 503) သည် Anthropic ပုံစံဖြင့် လာပြီး 401၊ 413၊ 415 နှင့် 422 တို့သည် OpenAI ပုံစံဖြင့် လာသည်။ status code ကို အရင်ဖတ်ပြီး ပုံစံနှစ်ခုလုံးတွင် ပါသော error.type နှင့် error.message ကို ဖတ်ပါ။
{
"error": {
"type": "invalid_request_error",
"message": "tokenize is available for the hosted open models; unknown model: shannon-3"
}
} {
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "count_tokens is available for the hosted open models; unknown model: shannon-3"
}
}