Skip to content
Token counting

Token counting

Count the tokens of a text or of a whole request before you send it.

POST https://api.shannon-ai.com/v1/tokenize

POST https://api.shannon-ai.com/v1/messages/count_tokens

Both endpoints count with the tokenizer of the model you name, and no model runs. They cover the hosted open-weight models. /v1/tokenize takes a plain text or a Chat Completions conversation. /v1/messages/count_tokens takes a request in the Anthropic Messages format, which is the call the Anthropic SDK and Claude Code make.

Counting is free. A call needs your API key, takes nothing from your balance and does not appear in your usage log.

Count a text

Send model and text. The text is counted as it is, with no chat formatting around it.

import requests

response = requests.post(
    "https://api.shannon-ai.com/v1/tokenize",
    headers={"Authorization": "Bearer YOUR_API_KEY"},
    json={
        "model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
        "text": "Hello, world",
    },
)
print(response.json()["tokens"])
200 Reply
{
  "model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
  "tokens": 3
}

The numbers in the replies on this page are examples. The same text gives a different count on a different model.

Count a chat request

Send model and messages, with tools when the request has them, exactly as you would send them to /v1/chat/completions. The reply is the size of the whole input.

import requests

request = {
    "model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
    "messages": [
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "What is the weather in Paris?"},
    ],
    "tools": [
        {
            "type": "function",
            "function": {
                "name": "get_weather",
                "description": "Current weather for a city",
                "parameters": {
                    "type": "object",
                    "properties": {"city": {"type": "string"}},
                    "required": ["city"],
                },
            },
        }
    ],
}

response = requests.post(
    "https://api.shannon-ai.com/v1/tokenize",
    headers={"Authorization": "Bearer YOUR_API_KEY"},
    json=request,
)
print(response.json()["tokens"])
200 Reply
{
  "model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
  "tokens": 164
}

Fields of /v1/tokenize

Field Type Description
model string Required. A hosted open-weight model id. Upper and lower case are treated the same.
text string A text to count as it is, with no chat formatting. Up to 4,000,000 bytes. Send text or messages; when both are present, text is counted.
messages array Chat messages in the Chat Completions format. They are counted as the full input of a request: every message with the formatting the model's chat template puts around it.
tools array Tool definitions to include in the count. Used together with messages.

The reply is a JSON object with these fields:

Field Type Description
model string The model id the count was made for, in its published spelling.
tokens integer With text: the tokens of the text. With messages: the tokens of the whole input, images included.

Count a Messages request

Send the body you would send to /v1/messages: model, messages, and system and tools when you use them. The official Anthropic SDKs call this endpoint through messages.count_tokens.

import anthropic

client = anthropic.Anthropic(
    api_key="YOUR_API_KEY",
    base_url="https://api.shannon-ai.com",
)

count = client.messages.count_tokens(
    model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
    system="You are a concise assistant.",
    messages=[
        {"role": "user", "content": "Summarise the attached report."}
    ],
)
print(count.input_tokens)
200 Reply
{
  "input_tokens": 21
}

Fields of /v1/messages/count_tokens

Field Type Description
model string Required. A hosted open-weight model id.
messages array Required. Messages in the Anthropic Messages format. text, image, tool_use and tool_result blocks are counted.
system string | array The system prompt: a string or an array of text blocks.
tools array Tool definitions with name, description and input_schema.

Accepted for compatibility, with no effect on the count: tool_choice, max_tokens, temperature, top_p, stop_sequences, stream, thinking. You can pass the body of a real request unchanged.

The reply is a JSON object with these fields:

Field Type Description
input_tokens integer The tokens of the whole input: system prompt, messages, tools and images.

Supported models

Both endpoints count for the hosted open-weight models. GET /v1/models lists /v1/tokenize and /v1/messages/count_tokens in the endpoints of each model that supports them. Any other model value, the Shannon ids included, is answered with 400.

  • DeepSeek-V4-Pro-0813-3BIT-REAP
  • GLM-5.2-3BIT-REAP
  • Kimi-K3-3BIT-REAP
  • Nemotron3Ultra-3BIT-REAP
  • MiniMax-M3-3BIT-REAP
  • DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP
  • Kimi-K2.6-W4A16-AUTOROUND-REAP
  • Laguna-S-2.1-W4A16-AUTOROUND-REAP
  • inkling-W4A16-AUTOROUND-REAP
  • MiMo-V2.5-Pro-W8A16
  • MiMo-V2.5-W8A16
  • Hy3-W8A16

For a Shannon model, read the token counts from the usage object of a reply.

How the count is made

Each model is counted with its own tokenizer and its own chat template. No estimate from characters or words is used.

What is counted Rule
A text The tokens of the string as sent. An empty string counts 0.
Messages The messages and tools are laid out with the model's own chat template, up to the point where the reply begins, and that whole prompt is counted.
Roles system, user, assistant and tool messages are counted. developer is counted as system. A message with no content and no tool call adds nothing.
Tool calls and results Tool calls of earlier assistant turns and their results are part of the count, on both endpoints.
Images An image sent inside the body (base64 or a data: URL) adds one token per 28 × 28 pixel patch: ceil(width / 28) × ceil(height / 28). An image given as an http(s) URL is not downloaded by these endpoints and counts 1,024.

Example: a 1,024 × 768 pixel image counts ceil(1024 / 28) × ceil(768 / 28) = 37 × 28 = 1,036 tokens.

The count and what a request is charged

The count of a whole request is made the same way as the input count of a real request with the same model, messages and tools. A reply reports that number as usage.prompt_tokens on Chat Completions, as usage.input_tokens on Responses, and as usage.input_tokens plus usage.cache_read_input_tokens on Messages.

  • The count is the input before the cached-input discount. A real request may read part of that input from cache and bill that part at the cached rate. Prompt caching
  • An image given as an http(s) URL counts 1,024 here. A real request downloads the image and counts it from its size in pixels, so the two numbers can differ. Send the image as base64 to get the same number.
  • Output is not part of the count. The reply of a real request is billed as output tokens on top, reasoning included.
  • A text count has no chat formatting. Use it to measure a document or a prompt part, and the messages form to measure a request.

To turn a count into a cost, multiply it by the model's input price per 1M tokens. Models & pricing

Limits

Limit Value Above it
Length of text 4,000,000 bytes (UTF-8) 413 with the message text too long
Request body 32 MiB 413
Per request One text or one conversation Send one request per text to count several texts.

Counting calls are not counted toward the limit of 120 requests per minute. Limits and balance

Errors

Status Type Message When
400 invalid_request_error tokenize is available for the hosted open models; unknown model: <model> /v1/tokenize with a model that is not a hosted open-weight id.
400 invalid_request_error count_tokens is available for the hosted open models; unknown model: <model> /v1/messages/count_tokens with a model that is not a hosted open-weight id, or without model.
400 invalid_request_error send `text` or `messages` /v1/tokenize with neither text nor messages.
401 authentication_error Missing authentication / Invalid API key No key was sent, or the key is not valid.
413 invalid_request_error text too long text is longer than 4,000,000 bytes. A body over 32 MiB is also answered with 413.
415 invalid_request_error Expected request with `Content-Type: application/json` The request has no JSON content type.
422 invalid_request_error Failed to deserialize the JSON body into the target type: … A required field is missing (model on /v1/tokenize, messages on /v1/messages/count_tokens) or a field has the wrong type.
503 api_error token counting is temporarily unavailable for this model The count cannot be made for this model at the moment. Try again later.

/v1/tokenize returns errors in the OpenAI shape. On /v1/messages/count_tokens the errors of the endpoint itself (400 for the model, 503) come in the Anthropic shape, and 401, 413, 415 and 422 come in the OpenAI shape. Read the status code first, then error.type and error.message, which are present in both shapes.

400 /v1/tokenize
{
  "error": {
    "type": "invalid_request_error",
    "message": "tokenize is available for the hosted open models; unknown model: shannon-3"
  }
}
400 /v1/messages/count_tokens
{
  "type": "error",
  "error": {
    "type": "invalid_request_error",
    "message": "count_tokens is available for the hosted open models; unknown model: shannon-3"
  }
}