Tlola ho ea ho litaba
Caching ea Prompt

Caching ea Prompt

KE AUTOMATIC

Hosted open-weight models li cache-a prompt prefixes tse phetileng ka bo eona. Ha request e qala ka system prompt, tools le melaetsa e tšoanang le request ea morao-rao ho model e tšoanang, prefix eo e baloa ho tsoa ho cache ebile e betaloa ka 25% ea input price ea model. Ha ho na letho la ho enable, ebile cache writes ke mahala.

Mokhoa oa tšebetso

  • Prefix, ka tatso — Prompt e baloa ka tatso: system prompt, tool definitions, ebe ke melaetsa. Cache e lekisa ho tloha qalong ea sequence eo ho fihla ho token ea pele e fapalaneng.
  • Se queleloa hit — Request eo prompt ea eona e qalang ka content e tšoanang le request ea morao-rao — hangata ke turn ea feta ea conversation e tšoanang ka melaetsa e mecha e eketsitsoeng. Prefix e lekang ke cached input; tsohle ka morao oa eona ke regular input.
  • Botlenya ba cache — Cache e boloka prompt ka li-block tsa tokens tse 1,568, kahoo prompt e khutšoanyane ho feta tokens tse ka bang 1,500 ha e cache-oe. Palo ea cached karabong ke palo ea hau ea input e atolositsoeng ka karolo ea cached ea prompt, e fokotsoeng ho tse fokolang. Ha e ame e le palo e atolositsoeng ea boholo ba block.
  • Ntle le hit — Kopo eo qalo ea eona e seng ho cache e lefisoa ka input rate e tloaelehileng. Ha ho nako ea bophelo e phatlalalitsoeng bakeng sa li-prompt tse cached mme hit ha e tiisetsoe: bala usage ho bona seo kopo e se nkileng ho cache.
  • Ha ho switch — Kopo ha e kenelle, mme ha ho field e tima caching.
  • Models lifelesso — Hosted open-weight id e mong’a. GET /v1/models e rapela capabilities.prompt_caching: true le pricing.cached_input_per_million_usd bakeng sa tsona. Shannon models li betala rate e le nngwe.

Bona cache hit karabong

Romela likopo tse peli tse qalang ka system prompt e telele e tšoanang mme u gatisetse usage ea e 'ngoe le e 'ngoe. Nomoro ea pele ke input ea kopo, ea bobeli ke karolo ea eona e baloang ho cache.

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")

handbook = open("handbook.txt").read()  # a long text that stays the same


def ask(question):
    response = client.chat.completions.create(
        model="Kimi-K3-3BIT-REAP",
        messages=[
            {"role": "system", "content": handbook},
            {"role": "user", "content": question},
        ],
    )
    usage = response.usage
    print(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens)


ask("What is the refund policy?")
ask("Who approves travel?")  # same start: read the second number

Theko

Cached input tokens li betaloa ka 25% ea input rate ea model, e kokotloang ho $0.001 per 1M. Ho ngola ho cache ha ho na lebeletso ea theko, ebile output e betaloa ka mokhoa oa tlhoahlo. Cached rate ea id ka’ng e teng tabuleng ea Models & pricing. Li-model le litheko

Input ea kopo e lefisoa e le (input − cached) × input rate + cached × cached rate. Palo ea cached ha e ka ke ea feta palo ea input.

Model Input / 1M Input e cached / 1M
DeepSeek-V4-Pro-0813-3BIT-REAP $1.95 $0.488
GLM-5.2-3BIT-REAP $0.73 $0.183
Kimi-K3-3BIT-REAP $3.83 $0.958
Nemotron3Ultra-3BIT-REAP $0.75 $0.188
MiniMax-M3-3BIT-REAP $0.50 $0.125
DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP $0.50 $0.125
Kimi-K2.6-W4A16-AUTOROUND-REAP $0.78 $0.195
Laguna-S-2.1-W4A16-AUTOROUND-REAP $0.50 $0.125
inkling-W4A16-AUTOROUND-REAP $1.42 $0.355
MiMo-V2.5-Pro-W8A16 $0.50 $0.125
MiMo-V2.5-W8A16 $0.50 $0.125
Hy3-W8A16 $0.50 $0.125

Usage log e thathamisa input e cached ea kopo e 'ngoe le e 'ngoe. Tokens tse lefisitsoeng le theko ea eona li se li kenyelletsa cached rate. Keys & usage

Fields tsa tšebeliso

Endpoint Input e cached Reasoning
/v1/chat/completions usage.prompt_tokens_details.cached_tokens — karolo ea prompt_tokens usage.completion_tokens_details.reasoning_tokens — karolo ea completion_tokens
/v1/responses usage.input_tokens_details.cached_tokens — karolo ea input_tokens usage.output_tokens_details.reasoning_tokens — karolo ea output_tokens
/v1/messages usage.cache_read_input_tokens — e rapeloa ka setala: input_tokens ke karolo e sa cached-oeng; cache_creation_input_tokens ke 0 kamehla thinking e baloa ho output_tokens
{
  "usage": {
    "prompt_tokens": 20000,
    "completion_tokens": 812,
    "total_tokens": 20812,
    "prompt_tokens_details": {
      "cached_tokens": 18000
    },
    "completion_tokens_details": {
      "reasoning_tokens": 604
    }
  }
}

Karabo ea stream e nka li-field tse tšoanang usage ea eona ea ho qetela. Ha u hloke ho e kopa:

Endpoint Moo usage e fihlang teng
/v1/chat/completions usage ho chunk ea ho qetela pele ho data: [DONE]. E romeloa ho stream e 'ngoe le e 'ngoe.
/v1/responses response.usage ea event ea response.completed.
/v1/messages usage ea event ea message_delta. usage ea message_start e na le li-zero.

Ho fumana cache hits tse ngata

  • Boloka system prompt le tool definitions li tšoana byte-ka-byte pakeng tsa likhoapolo. Beha values tsa call ka’ng joalo ka timestamps kapa request ids qetellong ea melaetsa e morao, eseng ho system prompt.
  • Eketsa feela ho history. Ho fetola, ho fokotsa kapa ho summarize turns tsa pele ho fetola prefix, ebile tsohle ka morao oa pheto ea pele li betaloa e le regular input.
  • u se fetole tatso ea tools, melaetsa kapa content blocks pakeng tsa likhoapolo, ebile serialise JSON (tool schemas, tool arguments le results) ka mokhoa o tšoanang kamehla.
  • Lula ho id e le 'ngoe ea model moqoqong, mme u romele kopo e latelang kapele ka mor'a e fetileng.

API e boloka qalo ea moqoqo e tsitsitse maemong ana:

  • Molaetsa oa system kapa developer o romeloang hamorao moqoqong o lula sebakeng sa oona. Ha o fetole qalo ea prompt, kahoo liphetoho tse o fetileng li lula li cached.
  • Li-argument tsa likopo tsa tool liphetohong tse fetileng tsa assistant li bapisoa ka boleng. Tatellano ea li-key le sebaka sa JSON eo ha li bohlokoa.
  • Li-endpoint tse tharo li bala moqoqo ka tsela e tšoanang. Moqoqo o tsoelang pele ho endpoint e 'ngoe o boloka prefix ea oona e arolelanoang ha litaba li tšoana.

Fields tsa kopo

prompt_cache_key (Chat Completions le Responses) le cache_control ho Message content blocks li ammehiloe, kahoo code ea client e seng elahlehile e sebetsa ntle le liphetwo. Ha ho na le eona ea hlokang: caching ke ea automatiki ’me e sebetsa ka tsela e tšoanang ntle le tsona.

Field E rometsoe ho Ke eng
prompt_cache_key /v1/chat/completions, /v1/responses Cache routing key ea OpenAI API.
cache_control /v1/messages Cache breakpoint ho content block, block ea system kapa molaetsa oa Anthropic API.
stream_options /v1/chat/completions include_usage e kopa OpenAI API usage ho stream. Mona stream e 'ngoe le e 'ngoe e fela ka usage.

Ho bala tokens

Li-endpoint tse peli tsa mahala, POST /v1/tokenize le POST /v1/messages/count_tokens, li bala tokens tsa mongolo kapa tsa kopo eohle bakeng sa li-model tse hosted tsa open-weight pele u e romela. Li na le leqephe la tsona: Ho bala tokens