Skip to content
Prompt caching

Prompt caching

AUTOMATIC

The hosted open-weight models cache repeated prompt prefixes automatically. When a request starts with the same system prompt, tools and earlier messages as a recent request on the same model, that shared prefix is read from cache and billed at 25% of the model's input price. There is nothing to enable, and cache writes are free.

How it works

  • Prefix, in order — The prompt is read in order: system prompt, tool definitions, then the messages. The cache matches from the start of that sequence up to the first token that differs.
  • What counts as a hit — A request whose prompt begins with the same content as a recent request — typically the previous turn of the same conversation with new messages appended. The matching prefix is cached input; everything after it is regular input.
  • Granularity — The cache holds a prompt in blocks of 1,568 tokens, so a prompt shorter than about 1,500 tokens is not cached. The cached count in a reply is your input count multiplied by the cached share of the prompt, rounded down. It is not necessarily a multiple of the block size.
  • Without a hit — A request whose start is not in the cache is billed at the regular input rate. No lifetime is published for cached prompts and a hit is not guaranteed: read usage to see what a request took from the cache.
  • No switch — A request does not opt in, and no field turns caching off.
  • Which models — Every hosted open-weight id. GET /v1/models reports capabilities.prompt_caching: true and pricing.cached_input_per_million_usd for them. Shannon models bill one flat rate.

See a cache hit in a reply

Send two requests that start with the same long system prompt and print the usage of each. The first number is the input of the request, the second is the part of it that was read from cache.

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")

handbook = open("handbook.txt").read()  # a long text that stays the same


def ask(question):
    response = client.chat.completions.create(
        model="Kimi-K3-3BIT-REAP",
        messages=[
            {"role": "system", "content": handbook},
            {"role": "user", "content": question},
        ],
    )
    usage = response.usage
    print(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens)


ask("What is the refund policy?")
ask("Who approves travel?")  # same start: read the second number

Pricing

Cached input tokens bill at 25% of the model's input rate, rounded to $0.001 per 1M. Writing to the cache costs nothing extra, and output is billed as usual. Each id's cached rate is in the Models & pricing table. Models & pricing

The input of a call is charged as (input − cached) × input rate + cached × cached rate. The cached count is never larger than the input count.

Model Input / 1M Cached input / 1M
DeepSeek-V4-Pro-0813-3BIT-REAP $1.95 $0.488
GLM-5.2-3BIT-REAP $0.73 $0.183
Kimi-K3-3BIT-REAP $3.83 $0.958
Nemotron3Ultra-3BIT-REAP $0.75 $0.188
MiniMax-M3-3BIT-REAP $0.50 $0.125
DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP $0.50 $0.125
Kimi-K2.6-W4A16-AUTOROUND-REAP $0.78 $0.195
Laguna-S-2.1-W4A16-AUTOROUND-REAP $0.50 $0.125
inkling-W4A16-AUTOROUND-REAP $1.42 $0.355
MiMo-V2.5-Pro-W8A16 $0.50 $0.125
MiMo-V2.5-W8A16 $0.50 $0.125
Hy3-W8A16 $0.50 $0.125

The usage log lists the cached input of each call. Its billed tokens and cost already include the cached rate. Keys & usage

Usage fields

Endpoint Cached input Reasoning
/v1/chat/completions usage.prompt_tokens_details.cached_tokens — part of prompt_tokens usage.completion_tokens_details.reasoning_tokens — part of completion_tokens
/v1/responses usage.input_tokens_details.cached_tokens — part of input_tokens usage.output_tokens_details.reasoning_tokens — part of output_tokens
/v1/messages usage.cache_read_input_tokens — reported apart: input_tokens is the uncached part; cache_creation_input_tokens is always 0 thinking is counted in output_tokens
{
  "usage": {
    "prompt_tokens": 20000,
    "completion_tokens": 812,
    "total_tokens": 20812,
    "prompt_tokens_details": {
      "cached_tokens": 18000
    },
    "completion_tokens_details": {
      "reasoning_tokens": 604
    }
  }
}

A streamed reply carries the same fields in its final usage. You do not have to ask for it:

Endpoint Where the usage arrives
/v1/chat/completions usage on the last chunk before data: [DONE]. It is sent on every stream.
/v1/responses response.usage of the response.completed event.
/v1/messages usage of the message_delta event. The usage of message_start holds zeros.

Getting more cache hits

  • Keep the system prompt and tool definitions byte-for-byte stable across calls. Put per-call values such as timestamps or request ids at the end of the latest message, not in the system prompt.
  • Only append to the history. Editing, trimming or summarising earlier turns changes the prefix, and everything after the first change is billed as regular input.
  • Do not reorder tools, messages or content blocks between calls, and serialise JSON (tool schemas, tool arguments and results) the same way every time.
  • Stay on one model id for a conversation, and send the follow-up call soon after the one before it.

The API keeps the start of a conversation stable in these cases:

  • A system or developer message sent later in a conversation stays at its place. It does not change the start of the prompt, so the turns before it stay cached.
  • The arguments of tool calls in earlier assistant turns are compared by value. Key order and spacing of that JSON do not matter.
  • The three endpoints read a conversation the same way. A conversation continued on another endpoint keeps its shared prefix when the content is the same.

Request fields

prompt_cache_key (Chat Completions and Responses) and cache_control on Messages content blocks are accepted, so existing client code runs unchanged. Neither is required: caching is automatic and works the same without them.

Field Sent to What it is
prompt_cache_key /v1/chat/completions, /v1/responses A cache routing key of the OpenAI API.
cache_control /v1/messages A cache breakpoint on a content block, a system block or a message of the Anthropic API.
stream_options /v1/chat/completions include_usage asks the OpenAI API for usage on a stream. Here every stream ends with usage.

Counting tokens

Two free endpoints, POST /v1/tokenize and POST /v1/messages/count_tokens, count the tokens of a text or of a whole request for the hosted open-weight models before you send it. They have their own page: Token counting