Token எண்ணிக்கை
ஒரு உரையின் அல்லது முழுக் கோரிக்கையின் tokens-ஐ அனுப்பும் முன் எண்ணிப் பாருங்கள்.
POST https://api.shannon-ai.com/v1/tokenize
POST https://api.shannon-ai.com/v1/messages/count_tokens
இரு endpoints-ம் நீங்கள் பெயரிடும் மாடலின் tokenizer கொண்டு எண்ணுகின்றன; எந்த மாடலும் இயக்கப்படுவதில்லை. அவை ஹோஸ்ட் செய்யப்பட்ட open-weight மாடல்களுக்கானவை. /v1/tokenize சாதாரண உரை அல்லது Chat Completions உரையாடலை ஏற்கும். /v1/messages/count_tokens Anthropic Messages வடிவ கோரிக்கையை ஏற்கும்; Anthropic SDK-யும் Claude Code-ம் செய்யும் அழைப்பு இதுவே.
எண்ணுவது இலவசம். அழைப்புக்கு உங்கள் API key தேவை; அது உங்கள் balance-ல் இருந்து எதையும் எடுக்காது, உங்கள் usage log-லும் தோன்றாது.
உரையை எண்ணுதல்
model மற்றும் text-ஐ அனுப்பவும். உரை அப்படியே, அதைச் சுற்றி chat வடிவமைப்பு இல்லாமல் எண்ணப்படும்.
import requests
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world",
},
)
print(response.json()["tokens"]) const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
text: "Hello, world",
}),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world"
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 3
} இந்தப் பக்கத்தில் உள்ள பதில்களின் எண்கள் எடுத்துக்காட்டுகள். அதே உரைக்கு வேறு மாடலில் வேறு எண்ணிக்கை கிடைக்கும்.
chat கோரிக்கையை எண்ணுதல்
model மற்றும் messages-ஐ அனுப்பவும்; கோரிக்கையில் இருந்தால் tools-ஐயும், /v1/chat/completions-க்கு அனுப்புவது போலவே. பதில் முழு உள்ளீட்டின் அளவு.
import requests
request = {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"},
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
],
}
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json=request,
)
print(response.json()["tokens"]) const request = {
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
messages: [
{ role: "system", content: "You are a concise assistant." },
{ role: "user", content: "What is the weather in Paris?" },
],
tools: [
{
type: "function",
function: {
name: "get_weather",
description: "Current weather for a city",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
},
},
},
],
};
const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify(request),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}
]
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 164
} /v1/tokenize-ன் புலங்கள்
| புலம் | வகை | விளக்கம் |
|---|---|---|
model | string | கட்டாயம். ஹோஸ்ட் செய்யப்பட்ட open-weight மாடல் id. பெரிய மற்றும் சிறிய எழுத்துகள் ஒன்றாகவே கருதப்படும். |
text | string | chat வடிவமைப்பு இல்லாமல், அப்படியே எண்ணப்பட வேண்டிய உரை. 4,000,000 பைட்டுகள் வரை. text அல்லது messages-ஐ அனுப்பவும்; இரண்டும் இருந்தால் text எண்ணப்படும். |
messages | array | Chat Completions வடிவில் chat messages. அவை கோரிக்கையின் முழு உள்ளீடாக எண்ணப்படும்: மாடலின் chat template ஒவ்வொரு செய்தியையும் சுற்றி வைக்கும் வடிவமைப்புடன். |
tools | array | எண்ணிக்கையில் சேர்க்க வேண்டிய tool வரையறைகள். messages உடன் சேர்த்துப் பயன்படுத்தவும். |
பதில் இந்தப் புலங்கள் கொண்ட JSON object:
| புலம் | வகை | விளக்கம் |
|---|---|---|
model | string | எண்ணிக்கை செய்யப்பட்ட மாடல் id, வெளியிடப்பட்ட எழுத்துவடிவில். |
tokens | integer | text உடன்: உரையின் tokens. messages உடன்: படங்கள் உட்பட முழு உள்ளீட்டின் tokens. |
Messages கோரிக்கையை எண்ணுதல்
/v1/messages-க்கு நீங்கள் அனுப்பும் அதே body-ஐ அனுப்பவும்: model, messages, மேலும் பயன்படுத்தினால் system மற்றும் tools. அதிகாரப்பூர்வ Anthropic SDK-கள் இந்த endpoint-ஐ messages.count_tokens வழியாக அழைக்கின்றன.
import anthropic
client = anthropic.Anthropic(
api_key="YOUR_API_KEY",
base_url="https://api.shannon-ai.com",
)
count = client.messages.count_tokens(
model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system="You are a concise assistant.",
messages=[
{"role": "user", "content": "Summarise the attached report."}
],
)
print(count.input_tokens) import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
apiKey: "YOUR_API_KEY",
baseURL: "https://api.shannon-ai.com",
});
const count = await client.messages.countTokens({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system: "You are a concise assistant.",
messages: [
{ role: "user", content: "Summarise the attached report." },
],
});
console.log(count.input_tokens); curl https://api.shannon-ai.com/v1/messages/count_tokens \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"system": "You are a concise assistant.",
"messages": [
{"role": "user", "content": "Summarise the attached report."}
]
}' {
"input_tokens": 21
} /v1/messages/count_tokens endpoint-ன் புலங்கள் பற்றிய விவரம்
| புலம் | வகை | விளக்கம் |
|---|---|---|
model | string | கட்டாயம். ஹோஸ்ட் செய்யப்பட்ட open-weight மாடல் id. |
messages | array | கட்டாயம். Anthropic Messages வடிவில் messages. text, image, tool_use மற்றும் tool_result blocks எண்ணப்படும். |
system | string | array | system prompt: ஒரு string அல்லது உரைத் தொகுதிகளின் (text blocks) அணி. |
tools | array | name, description மற்றும் input_schema கொண்ட tool வரையறைகள். |
இணக்கத்தன்மைக்காக ஏற்றுக்கொள்ளப்படும், எண்ணிக்கையில் எந்த விளைவும் இல்லாமல்: tool_choice, max_tokens, temperature, top_p, stop_sequences, stream, thinking. உண்மையான கோரிக்கையின் body-ஐ மாற்றமின்றி அனுப்பலாம்.
பதில் இந்தப் புலங்கள் கொண்ட JSON object:
| புலம் | வகை | விளக்கம் |
|---|---|---|
input_tokens | integer | முழு உள்ளீட்டின் tokens: system prompt, messages, tools மற்றும் படங்கள். |
ஆதரிக்கப்படும் மாடல்கள்
இரு endpoints-ம் ஹோஸ்ட் செய்யப்பட்ட open-weight மாடல்களுக்கு எண்ணும். GET /v1/models ஒவ்வொரு ஆதரவு மாடலின் endpoints-ல் /v1/tokenize மற்றும் /v1/messages/count_tokens-ஐப் பட்டியலிடுகிறது. வேறு எந்த model மதிப்புக்கும், Shannon ids உட்பட, 400 பதில் வரும்.
DeepSeek-V4-Pro-0813-3BIT-REAPGLM-5.2-3BIT-REAPKimi-K3-3BIT-REAPNemotron3Ultra-3BIT-REAPMiniMax-M3-3BIT-REAPDeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAPKimi-K2.6-W4A16-AUTOROUND-REAPLaguna-S-2.1-W4A16-AUTOROUND-REAPinkling-W4A16-AUTOROUND-REAPMiMo-V2.5-Pro-W8A16MiMo-V2.5-W8A16Hy3-W8A16
Shannon மாடலுக்கு, tokens எண்ணிக்கையைப் பதிலின் usage object-ல் இருந்து படிக்கவும்.
எண்ணிக்கை எவ்வாறு செய்யப்படுகிறது
ஒவ்வொரு மாடலும் அதன் சொந்த tokenizer மற்றும் சொந்த chat template கொண்டு எண்ணப்படுகிறது. எழுத்துகள் அல்லது சொற்களின் அடிப்படையில் மதிப்பீடு எதுவும் பயன்படுத்தப்படுவதில்லை.
| எது எண்ணப்படுகிறது | விதி |
|---|---|
| ஒரு உரை | அனுப்பியபடியே string-ன் tokens. காலி string 0 ஆகக் கணக்கிடப்படும். |
| செய்திகள் | messages மற்றும் tools மாடலின் சொந்த chat template படி, பதில் தொடங்கும் புள்ளி வரை அமைக்கப்படும்; அந்த முழு prompt-ம் எண்ணப்படும். |
| பாத்திரங்கள் (Roles) | system, user, assistant மற்றும் tool செய்திகள் எண்ணப்படும். developer என்பது system ஆகக் கணக்கிடப்படும். உள்ளடக்கமும் tool அழைப்பும் இல்லாத செய்தி எதையும் சேர்க்காது. |
| Tool அழைப்புகளும் முடிவுகளும் | முந்தைய assistant சுற்றுகளின் tool அழைப்புகளும் அவற்றின் முடிவுகளும் இரு endpoints-லும் எண்ணிக்கையின் பகுதியாகும். |
| படங்கள் | body-க்குள் அனுப்பப்படும் படம் (base64 அல்லது data: URL) ஒவ்வொரு 28 × 28 பிக்சல் பகுதிக்கும் ஒரு token சேர்க்கும்: ceil(width / 28) × ceil(height / 28). http(s) URL ஆகக் கொடுக்கப்பட்ட படத்தை இந்த endpoints பதிவிறக்குவதில்லை; அது 1,024 ஆகக் கணக்கிடப்படும். |
எடுத்துக்காட்டு: 1,024 × 768 பிக்சல் படம் ceil(1024 / 28) × ceil(768 / 28) = 37 × 28 = 1,036 tokens ஆகக் கணக்கிடப்படும்.
எண்ணிக்கையும் கோரிக்கைக்கான கட்டணமும்
முழுக் கோரிக்கையின் எண்ணிக்கை, அதே மாடல், messages மற்றும் tools கொண்ட உண்மையான கோரிக்கையின் உள்ளீட்டு எண்ணிக்கை போலவே கணக்கிடப்படுகிறது. பதில் அந்த எண்ணை Chat Completions-ல் usage.prompt_tokens ஆகவும், Responses-ல் usage.input_tokens ஆகவும், Messages-ல் usage.input_tokens மற்றும் usage.cache_read_input_tokens கூட்டுத்தொகையாகவும் தெரிவிக்கிறது.
- எண்ணிக்கை என்பது cached-input தள்ளுபடிக்கு முந்தைய உள்ளீடு. உண்மையான கோரிக்கை அந்த உள்ளீட்டின் ஒரு பகுதியை cache-ல் இருந்து படித்து, அந்தப் பகுதிக்கு cached விலையில் கட்டணம் விதிக்கலாம். பிராம்ப்ட் கேச்சிங்
http(s)URL ஆகக் கொடுக்கப்பட்ட படம் இங்கு 1,024 ஆகக் கணக்கிடப்படும். உண்மையான கோரிக்கை படத்தைப் பதிவிறக்கி, அதன் பிக்சல் அளவிலிருந்து கணக்கிடும்; எனவே இரு எண்களும் வேறுபடலாம். அதே எண்ணைப் பெற படத்தை base64 ஆக அனுப்பவும்.- வெளியீடு எண்ணிக்கையின் பகுதி அல்ல. உண்மையான கோரிக்கையின் பதிலுக்கு, reasoning உட்பட, வெளியீட்டு tokens ஆகக் கூடுதலாகக் கட்டணம் விதிக்கப்படும்.
textஎண்ணிக்கையில் chat வடிவமைப்பு இல்லை. ஆவணம் அல்லது prompt-ன் ஒரு பகுதியை அளக்க இதையும், கோரிக்கையை அளக்கmessagesவடிவத்தையும் பயன்படுத்தவும்.
எண்ணிக்கையைச் செலவாக மாற்ற, அதை மாடலின் 1M tokens-க்கான உள்ளீட்டு விலையால் பெருக்கவும். மாடல்கள் & விலை
வரம்புகள்
| வரம்பு | மதிப்பு | அதற்கு மேல் |
|---|---|---|
text-ன் நீளம் | 4,000,000 பைட்டுகள் (UTF-8) | text too long என்ற செய்தியுடன் 413 |
| கோரிக்கை body | 32 MiB | 413 |
| ஒரு கோரிக்கைக்கு | ஒரு உரை அல்லது ஒரு உரையாடல் | பல உரைகளை எண்ண, ஒவ்வொரு உரைக்கும் தனிக் கோரிக்கை அனுப்பவும். |
எண்ணும் அழைப்புகள் நிமிடத்திற்கு 120 கோரிக்கைகள் என்ற வரம்பில் சேர்க்கப்படுவதில்லை. வரம்புகளும் இருப்பும்
பிழைகள்
| நிலை | வகை | செய்தி | எப்போது |
|---|---|---|---|
400 | invalid_request_error | tokenize is available for the hosted open models; unknown model: <model> | ஹோஸ்ட் செய்யப்பட்ட open-weight id அல்லாத model உடன் /v1/tokenize. |
400 | invalid_request_error | count_tokens is available for the hosted open models; unknown model: <model> | ஹோஸ்ட் செய்யப்பட்ட open-weight id அல்லாத model உடன், அல்லது model இல்லாமல் /v1/messages/count_tokens. |
400 | invalid_request_error | send `text` or `messages` | text அல்லது messages எதுவும் இல்லாமல் /v1/tokenize. |
401 | authentication_error | Missing authentication / Invalid API key | key அனுப்பப்படவில்லை, அல்லது key செல்லுபடியாகாது. |
413 | invalid_request_error | text too long | text 4,000,000 பைட்டுகளை விட நீளமானது. 32 MiB-ஐ மீறும் body-க்கும் 413 பதில் வரும். |
415 | invalid_request_error | Expected request with `Content-Type: application/json` | கோரிக்கையில் JSON content type இல்லை. |
422 | invalid_request_error | Failed to deserialize the JSON body into the target type: … | கட்டாயப் புலம் இல்லை (/v1/tokenize-ல் model, /v1/messages/count_tokens-ல் messages) அல்லது புலத்தின் வகை தவறானது. |
503 | api_error | token counting is temporarily unavailable for this model | இந்த மாடலுக்கு இப்போது எண்ணிக்கை செய்ய இயலாது. பின்னர் மீண்டும் முயற்சிக்கவும். |
/v1/tokenize பிழைகளை OpenAI வடிவில் தருகிறது. /v1/messages/count_tokens-ல் endpoint-ன் சொந்தப் பிழைகள் (மாடலுக்கு 400, 503) Anthropic வடிவில் வரும்; 401, 413, 415 மற்றும் 422 OpenAI வடிவில் வரும். முதலில் நிலைக் குறியீட்டைப் படிக்கவும், பிறகு இரு வடிவங்களிலும் உள்ள error.type மற்றும் error.message-ஐப் படிக்கவும்.
{
"error": {
"type": "invalid_request_error",
"message": "tokenize is available for the hosted open models; unknown model: shannon-3"
}
} {
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "count_tokens is available for the hosted open models; unknown model: shannon-3"
}
}