토큰 수 계산
텍스트나 전체 요청을 보내기 전에 토큰 수를 계산하세요.
POST https://api.shannon-ai.com/v1/tokenize
POST https://api.shannon-ai.com/v1/messages/count_tokens
두 엔드포인트 모두 지정한 모델의 토크나이저로 계산하며 모델은 실행되지 않습니다. 호스팅 오픈 웨이트 모델이 대상입니다. /v1/tokenize는 일반 텍스트 또는 Chat Completions 대화를 받습니다. /v1/messages/count_tokens는 Anthropic Messages 형식의 요청을 받으며, Anthropic SDK와 Claude Code가 보내는 호출이 바로 이것입니다.
토큰 수 계산은 무료입니다. 호출에는 API 키가 필요하지만 잔액에서 아무것도 차감되지 않고 사용량 기록에도 나타나지 않습니다.
텍스트의 토큰 수 계산
model과 text를 보내세요. 텍스트는 채팅 서식 없이 있는 그대로 계산됩니다.
import requests
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world",
},
)
print(response.json()["tokens"]) const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
text: "Hello, world",
}),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"text": "Hello, world"
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 3
} 이 페이지의 응답에 있는 숫자는 예시입니다. 같은 텍스트라도 모델이 다르면 수가 달라집니다.
채팅 요청의 토큰 수 계산
/v1/chat/completions에 보낼 때와 똑같이 model과 messages를 보내고, 요청에 도구가 있으면 tools도 보내세요. 응답은 전체 입력의 크기입니다.
import requests
request = {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"},
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
],
}
response = requests.post(
"https://api.shannon-ai.com/v1/tokenize",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json=request,
)
print(response.json()["tokens"]) const request = {
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
messages: [
{ role: "system", content: "You are a concise assistant." },
{ role: "user", content: "What is the weather in Paris?" },
],
tools: [
{
type: "function",
function: {
name: "get_weather",
description: "Current weather for a city",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
},
},
},
],
};
const response = await fetch("https://api.shannon-ai.com/v1/tokenize", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify(request),
});
const { tokens } = await response.json();
console.log(tokens); curl https://api.shannon-ai.com/v1/tokenize \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the weather in Paris?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}
]
}' {
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"tokens": 164
} /v1/tokenize의 필드 목록과 설명
| 필드 | 유형 | 설명 |
|---|---|---|
model | string | 필수. 호스팅 오픈 웨이트 모델 id입니다. 대소문자는 같게 처리됩니다. |
text | string | 채팅 서식 없이 그대로 계산할 텍스트입니다. 최대 4,000,000바이트입니다. text 또는 messages를 보내세요. 둘 다 있으면 text가 계산됩니다. |
messages | array | Chat Completions 형식의 채팅 메시지입니다. 요청의 전체 입력으로 계산되며, 모든 메시지는 모델의 채팅 템플릿이 주변에 붙이는 서식과 함께 계산됩니다. |
tools | array | 계산에 포함할 도구 정의입니다. messages와 함께 사용합니다. |
응답은 다음 필드가 있는 JSON 객체입니다.
| 필드 | 유형 | 설명 |
|---|---|---|
model | string | 토큰 수를 계산한 모델 id이며 공개된 표기로 반환됩니다. |
tokens | integer | text를 보내면 텍스트의 토큰 수입니다. messages를 보내면 이미지를 포함한 전체 입력의 토큰 수입니다. |
Messages 요청의 토큰 수 계산
/v1/messages에 보낼 본문을 그대로 보내세요: model, messages, 그리고 사용한다면 system과 tools. 공식 Anthropic SDK는 messages.count_tokens로 이 엔드포인트를 호출합니다.
import anthropic
client = anthropic.Anthropic(
api_key="YOUR_API_KEY",
base_url="https://api.shannon-ai.com",
)
count = client.messages.count_tokens(
model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system="You are a concise assistant.",
messages=[
{"role": "user", "content": "Summarise the attached report."}
],
)
print(count.input_tokens) import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
apiKey: "YOUR_API_KEY",
baseURL: "https://api.shannon-ai.com",
});
const count = await client.messages.countTokens({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
system: "You are a concise assistant.",
messages: [
{ role: "user", content: "Summarise the attached report." },
],
});
console.log(count.input_tokens); curl https://api.shannon-ai.com/v1/messages/count_tokens \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"system": "You are a concise assistant.",
"messages": [
{"role": "user", "content": "Summarise the attached report."}
]
}' {
"input_tokens": 21
} /v1/messages/count_tokens 요청에 들어 있는 필드 목록과 설명
| 필드 | 유형 | 설명 |
|---|---|---|
model | string | 필수. 호스팅 오픈 웨이트 모델 id입니다. |
messages | array | 필수. Anthropic Messages 형식의 메시지입니다. text, image, tool_use, tool_result 블록이 계산됩니다. |
system | string | array | 시스템 프롬프트: 문자열 또는 텍스트 블록의 배열입니다. |
tools | array | name, description, input_schema가 있는 도구 정의입니다. |
호환성을 위해 허용되며 계산에는 영향을 주지 않습니다: tool_choice, max_tokens, temperature, top_p, stop_sequences, stream, thinking. 실제 요청의 본문을 변경 없이 전달해도 됩니다.
응답은 다음 필드가 있는 JSON 객체입니다.
| 필드 | 유형 | 설명 |
|---|---|---|
input_tokens | integer | 전체 입력(시스템 프롬프트, 메시지, 도구, 이미지)의 토큰입니다. |
지원하는 모델
두 엔드포인트 모두 호스팅 오픈 웨이트 모델에 대해 계산합니다. GET /v1/models는 이를 지원하는 각 모델의 endpoints에 /v1/tokenize와 /v1/messages/count_tokens를 나열합니다. Shannon id를 포함한 다른 모든 model 값은 400으로 응답합니다.
DeepSeek-V4-Pro-0813-3BIT-REAPGLM-5.2-3BIT-REAPKimi-K3-3BIT-REAPNemotron3Ultra-3BIT-REAPMiniMax-M3-3BIT-REAPDeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAPKimi-K2.6-W4A16-AUTOROUND-REAPLaguna-S-2.1-W4A16-AUTOROUND-REAPinkling-W4A16-AUTOROUND-REAPMiMo-V2.5-Pro-W8A16MiMo-V2.5-W8A16Hy3-W8A16
Shannon 모델은 응답의 usage 객체에서 토큰 수를 읽으세요.
토큰 수가 계산되는 방식
각 모델은 고유의 토크나이저와 고유의 채팅 템플릿으로 계산됩니다. 글자 수나 단어 수로 하는 추정은 사용하지 않습니다.
| 계산 대상 | 규칙 |
|---|---|
| 텍스트 | 보낸 그대로의 문자열 토큰입니다. 빈 문자열은 0으로 계산됩니다. |
| 메시지 | 메시지와 도구는 모델 고유의 채팅 템플릿으로 응답이 시작되는 지점까지 배치되며, 그 프롬프트 전체가 계산됩니다. |
| 역할 | system, user, assistant, tool 메시지가 계산됩니다. developer는 system으로 계산됩니다. 콘텐츠도 도구 호출도 없는 메시지는 아무것도 더하지 않습니다. |
| 도구 호출과 결과 | 이전 assistant 턴의 도구 호출과 그 결과는 두 엔드포인트 모두에서 계산에 포함됩니다. |
| 이미지 | 본문 안에 담아 보낸 이미지(base64 또는 data: URL)는 28 × 28 픽셀 패치마다 토큰 하나를 더합니다: ceil(width / 28) × ceil(height / 28). http(s) URL로 지정한 이미지는 이 엔드포인트가 내려받지 않으며 1,024로 계산됩니다. |
예: 1,024 × 768 픽셀 이미지는 ceil(1024 / 28) × ceil(768 / 28) = 37 × 28 = 1,036 토큰으로 계산됩니다.
토큰 수와 요청에 청구되는 금액
전체 요청의 수는 같은 모델, 메시지, 도구로 보낸 실제 요청의 입력 수와 같은 방식으로 계산됩니다. 응답은 이 숫자를 Chat Completions에서는 usage.prompt_tokens, Responses에서는 usage.input_tokens, Messages에서는 usage.input_tokens와 usage.cache_read_input_tokens의 합으로 보고합니다.
- 이 수는 캐시된 입력 할인 전의 입력입니다. 실제 요청은 그 입력의 일부를 캐시에서 읽고 그 부분을 캐시 요율로 청구할 수 있습니다. 프롬프트 캐싱
http(s)URL로 지정한 이미지는 여기서 1,024로 계산됩니다. 실제 요청은 이미지를 내려받아 픽셀 크기로 계산하므로 두 숫자가 다를 수 있습니다. 같은 숫자를 얻으려면 이미지를 base64로 보내세요.- 출력은 이 수에 포함되지 않습니다. 실제 요청의 응답은 추론을 포함한 출력 토큰으로 추가 청구됩니다.
text수에는 채팅 서식이 없습니다. 문서나 프롬프트의 일부를 측정할 때 사용하고, 요청을 측정할 때는messages형태를 사용하세요.
수를 비용으로 바꾸려면 모델의 1M tokens당 입력 가격을 곱하세요. 모델 및 가격
한도
| 한도 | 값 | 초과 시 |
|---|---|---|
text의 길이 | 4,000,000바이트(UTF-8) | 메시지 text too long과 함께 413 |
| 요청 본문 | 32 MiB | 413 |
| 요청당 | 텍스트 하나 또는 대화 하나 | 여러 텍스트를 계산하려면 텍스트마다 요청을 하나씩 보내세요. |
토큰 수 계산 호출은 분당 120개 요청 한도에 포함되지 않습니다. 한도 및 잔액
오류
| 상태 | 타입 | 메시지 | 시점 |
|---|---|---|---|
400 | invalid_request_error | tokenize is available for the hosted open models; unknown model: <model> | 호스팅 오픈 웨이트 id가 아닌 model로 /v1/tokenize를 호출했습니다. |
400 | invalid_request_error | count_tokens is available for the hosted open models; unknown model: <model> | 호스팅 오픈 웨이트 id가 아닌 model로 /v1/messages/count_tokens를 호출했거나 model이 없습니다. |
400 | invalid_request_error | send `text` or `messages` | text도 messages도 없는 /v1/tokenize입니다. |
401 | authentication_error | Missing authentication / Invalid API key | 키를 보내지 않았거나 키가 유효하지 않습니다. |
413 | invalid_request_error | text too long | text가 4,000,000바이트보다 깁니다. 32 MiB를 넘는 본문도 413으로 응답합니다. |
415 | invalid_request_error | Expected request with `Content-Type: application/json` | 요청에 JSON 콘텐츠 타입이 없습니다. |
422 | invalid_request_error | Failed to deserialize the JSON body into the target type: … | 필수 필드가 없거나(/v1/tokenize의 model, /v1/messages/count_tokens의 messages) 필드의 타입이 잘못되었습니다. |
503 | api_error | token counting is temporarily unavailable for this model | 현재 이 모델은 수를 계산할 수 없습니다. 나중에 다시 시도하세요. |
/v1/tokenize는 OpenAI 형태로 오류를 반환합니다. /v1/messages/count_tokens에서는 엔드포인트 자체의 오류(모델에 대한 400, 503)는 Anthropic 형태로, 401, 413, 415, 422는 OpenAI 형태로 옵니다. 상태 코드를 먼저 읽고, 그다음 두 형태 모두에 있는 error.type과 error.message를 읽으세요.
{
"error": {
"type": "invalid_request_error",
"message": "tokenize is available for the hosted open models; unknown model: shannon-3"
}
} {
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "count_tokens is available for the hosted open models; unknown model: shannon-3"
}
}