Mandrosoa any amin'ny votoaty
Fitehirizana prompt

Fitehirizana prompt

AUTOMATIKA

Ny hosted open-weight models dia manao cache ny prompt prefixes miverimberina ho azy. Rehefa manomboka amin'ny system prompt, tools ary hafatra mitovy amin'ny request vao haingana amin'ny model iray ny request iray, dia vakiana avy amin'ny cache izany prefix iombonana izany ary aloa amin'ny 25% ny vidin'ny input an'ny model. Tsy misy zavatra tokony hampandehanana, ary maimaimpoana ny fanoratana cache.

Ny fomba fiasany

  • Prefix, araka ny filaharany — Vakiana araka ny filaharany ny prompt: system prompt, famaritana tool, avy eo ny hafatra. Ny cache dia mifanaraka manomboka amin'ny fiandohan'izany sequence izany hatramin'ny token voalohany izay tsy mitovy.
  • Inona no heverina ho 'hit' — Request izay manomboka amin'ny votoatiny mitovy amin'ny request vao haingana — matetika ny dingana teo aloha amin'ny resaka iray izay nampiana hafatra vaovao. Ny prefix mifanaraka dia input cached; ny ambiny rehetra aorian'izay dia input tsotra.
  • Granularité — Ny cache dia mitazona prompt amin'ny block 1,568 tokens, ka ny prompt fohy noho ny tokens 1,500 eo ho eo dia tsy cached. Ny isan'ny cached ao amin'ny valiny dia ny isan'ny input-nao ampitomboina amin'ny anjaran'ny prompt cached, ahena midina ho isa manontolo. Tsy voatery ho maromaro amin'ny haben'ny block izy.
  • Raha tsy misy hit — Ny fangatahana tsy ao anaty cache ny fiandohany dia lany amin'ny input rate mahazatra. Tsy misy fe-potoana navoaka ho an'ny prompt cached ary tsy azo antoka ny hit: vakio ny usage mba hahitana izay noraisin'ny fangatahana avy ao amin'ny cache.
  • Tsy misy switch — Tsy misy fangatahana mifidy ny caching, ary tsy misy saha manafoana azy.
  • Model inona — Ny id hosted open-weight rehetra. Ny GET /v1/models dia manome tatitra momba ny capabilities.prompt_caching: true sy ny pricing.cached_input_per_million_usd ho azy ireo. Ny Shannon models kosa dia mampiasa sarany raikitra iray.

Jereo ny cache hit ao amin'ny valiny

Alefaso fangatahana roa manomboka amin'ny system prompt lava mitovy ary asehoy ny usage tsirairay. Ny laharana voalohany dia ny input an'ny fangatahana, ny faharoa dia ny ampahany azo novakina avy ao amin'ny cache.

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")

handbook = open("handbook.txt").read()  # a long text that stays the same


def ask(question):
    response = client.chat.completions.create(
        model="Kimi-K3-3BIT-REAP",
        messages=[
            {"role": "system", "content": handbook},
            {"role": "user", "content": question},
        ],
    )
    usage = response.usage
    print(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens)


ask("What is the refund policy?")
ask("Who approves travel?")  # same start: read the second number

Vidiny

Ny cached input tokens dia aloa amin'ny 25% ny sarany miditra an'ny model, boriboriana ho $0.001 isaky ny 1M. Ny fanoratana cache dia tsy misy sarany fanampiny, ary ny output dia aloa araka ny mahazatra. Ny sarany cached an'ny id tsirairay dia ao amin'ny tabilao Models & pricing. Model & vidiny

Ny input amin'ny antso iray dia saina araka ny (input − cached) × input rate + cached × cached rate. Tsy lehibe kokoa noho ny isan'ny input mihitsy ny isan'ny cached.

Model Input / 1M Input cached / 1M
DeepSeek-V4-Pro-0813-3BIT-REAP $1.95 $0.488
GLM-5.2-3BIT-REAP $0.73 $0.183
Kimi-K3-3BIT-REAP $3.83 $0.958
Nemotron3Ultra-3BIT-REAP $0.75 $0.188
MiniMax-M3-3BIT-REAP $0.50 $0.125
DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP $0.50 $0.125
Kimi-K2.6-W4A16-AUTOROUND-REAP $0.78 $0.195
Laguna-S-2.1-W4A16-AUTOROUND-REAP $0.50 $0.125
inkling-W4A16-AUTOROUND-REAP $1.42 $0.355
MiMo-V2.5-Pro-W8A16 $0.50 $0.125
MiMo-V2.5-W8A16 $0.50 $0.125
Hy3-W8A16 $0.50 $0.125

Ny usage log dia mitanisa ny input cached isaky ny antso. Ny tokens nalaina sy ny vidiny dia efa ahitana ny cached rate. Keys & usage

Saha fampiasana

Endpoint Input cached Fandinihana
/v1/chat/completions usage.prompt_tokens_details.cached_tokens — ampahany amin'ny prompt_tokens usage.completion_tokens_details.reasoning_tokens — ampahany amin'ny completion_tokens
/v1/responses usage.input_tokens_details.cached_tokens — ampahany amin'ny input_tokens usage.output_tokens_details.reasoning_tokens — ampahany amin'ny output_tokens
/v1/messages usage.cache_read_input_tokens — tatitra mitokana: ny input_tokens no ampahany tsy cached; ny cache_creation_input_tokens dia 0 foana ny fisainana (thinking) dia isaina ao amin'ny output_tokens
{
  "usage": {
    "prompt_tokens": 20000,
    "completion_tokens": 812,
    "total_tokens": 20812,
    "prompt_tokens_details": {
      "cached_tokens": 18000
    },
    "completion_tokens_details": {
      "reasoning_tokens": 604
    }
  }
}

Ny valiny streamed dia mitondra saha mitovy ao amin'ny usage farany. Tsy ilaina angatahana izany:

Endpoint Aiza no tonga ny usage
/v1/chat/completions usage ao amin'ny chunk farany alohan'ny data: [DONE]. Alefa amin'ny stream rehetra izy.
/v1/responses response.usage an'ny event response.completed.
/v1/messages usage an'ny event message_delta. Ny usage an'ny message_start dia misy aotra.

Fomba hahazoana cache hits bebe kokoa

  • Tazony ho tsy miova tanteraka ny system prompt sy ny famaritana tool isaky ny call. Ataovy any amin'ny faran'ny hafatra farany ny sanda isaky ny call toy ny timestamps na request ids, fa tsy ao amin'ny system prompt.
  • Ampio fotsiny ny tantara (history). Ny fanovana, fanalana na famintanana ny dingana teo aloha dia manova ny prefix, ary ny zava-drehetra aorian'ny fanovana voalohany dia aloa amin'ny maha input tsotra azy.
  • Aza manova ny filaharan'ny tools, hafatra na content blocks eo anelanelan'ny calls, ary ampiasao ny fomba fanoratana JSON (tool schemas, tool arguments ary results) mitovy foana isaky ny mandefa.
  • Mijanòna amin'ny model id iray mandritra ny resaka, ary alefaso tsy ho ela ny antso manaraka aorian'ny teo aloha.

Ny API dia mitazona ny fiandohan'ny resaka tsy miova amin'ireto tranga ireto:

  • Ny hafatra system na developer alefa taty aoriana amin'ny resaka dia mijanona eo amin'ny toerany. Tsy manova ny fiandohan'ny prompt izy, ka mijanona cached ny dingana teo alohany.
  • Ny arguments an'ny tool calls amin'ny dingana assistant teo aloha dia ampitahaina araka ny sanda. Tsy mampaninona ny filaharan'ny key sy ny elanelana ao amin'io JSON io.
  • Ny endpoint telo dia mamaky resaka amin'ny fomba mitovy. Ny resaka tohizana amin'ny endpoint hafa dia mitazona ny prefix iombonana raha mitovy ny votoaty.

Sahan'ny fangatahana

Ekena ny prompt_cache_key (Chat Completions sy Responses) ary ny cache_control amin'ny vata vontoatin'ny Messages, koa tsy miova ny kaody client misy. Tsy voatery ny iray amin'ireo: mandeha ho azy ny caching ary miasa toy izany na tsy misy azy ireo aza.

Saha Alefa any amin'ny Inona izany
prompt_cache_key /v1/chat/completions, /v1/responses Cache routing key ao amin'ny API OpenAI.
cache_control /v1/messages Cache breakpoint amin'ny content block, block system na hafatra ao amin'ny API Anthropic.
stream_options /v1/chat/completions Ny include_usage dia mangataka amin'ny API OpenAI ny usage amin'ny stream. Eto dia mifarana miaraka amin'ny usage ny stream rehetra.

Fanisana tokens

Endpoint roa maimaim-poana, POST /v1/tokenize sy POST /v1/messages/count_tokens, no manisa ny tokens an'ny soratra na fangatahana manontolo ho an'ny model open-weight hosted alohan'ny handefasanao azy. Manana pejy manokana izy ireo: Fanisana token