Peke ki ngā ihirangi
Tūhono caching

Tūhono caching

AUNOA

Ko ngā hosted open-weight models e cache-tia ana ngā prefix prompt whakaherekaia a-automa tika. Ki te tīmataria he request ki te system prompt, ngā tools, me ngā karere rite ki tētahi request hou i te model kotahi, ka anau taua prefix mai i te cache, a ka ututia i te 25% o te utu input a te model. Kāore he mea hei whakahou, ā, he kore utu te tuhi ki te cache.

Me pēhea tana mahi

  • Prefix, i te huinga — Ka anau te prompt i te huinga: system prompt, ngā tohutohu tool, nā reira ngā karere. Ka rite te cache mai i te tīmatanga o taua huinga ki te token tuatahi e hē ana.
  • He aha te 'hit' — He request e tīmata ana tōna prompt ki te ihiraga rite ki tētahi request hou — ā-tikanga, ko te huringa mua o te kōrero kotahi me ngā karere hou i tāpiringia. Ko te prefix rite tēnei he input cached; ko-kete katoa whai ake he input noa.
  • Taumata taipitopito — Ka pupuri te cache i tētahi prompt i ngā poraka o te 1,568 tokens, nō reira kāore e cache te prompt iti iho i te 1,500 tokens tata. Ko te tatau cached i tētahi whakautu ko tō tatau input whakarea ki te wāhanga cached o te prompt, whakahekea. Kāore e kore he whakarea o te rahi poraka.
  • Kāore he hit — Ka utua te tono kāore tōna tīmatanga i te cache i te input rate auau. Kāore he wā ora e whakaputaina ana mō ngā prompt cached, ā, kāore te hit e whakapaingia: pānuitia te usage hei kite i tā te tono i tango mai i te cache.
  • Kāore he pana — Kāore te tono e whakaae, ā, kāore he āpure hei whakawetohia te caching.
  • No aupapa models — Katoa ngā hosted open-weight id. Ko te GET /v1/models e ripana ana i te capabilities.prompt_caching: true me te pricing.cached_input_per_million_usd mō rātou. Ko ngā Shannon models he reiti kotahi anau.

Kite i tētahi cache hit i tētahi whakautu

Tukuna e rua ngā tono e tīmata ana me te system prompt roa ōrite ka taka i te usage o ia tētahi. Ko te tau tuatahi te input o te tono, ko te tuarua te wāhanga i pānuitia mai i te cache.

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")

handbook = open("handbook.txt").read()  # a long text that stays the same


def ask(question):
    response = client.chat.completions.create(
        model="Kimi-K3-3BIT-REAP",
        messages=[
            {"role": "system", "content": handbook},
            {"role": "user", "content": question},
        ],
    )
    usage = response.usage
    print(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens)


ask("What is the refund policy?")
ask("Who approves travel?")  # same start: read the second number

Utu

Ka utu ngā cached input tokens i te 25% o te reiti input a te model, ka whakatau ki te $0.001 per 1M. Kāore he utu tari a te tuhi ki te cache, ā, ka utu te output hei tikanga noa. Kei te tēpu Models & pricing te reiti cached mō ia id. Ngā model me ngā utu

Ka utua te input o tētahi karanga hei (input − cached) × input rate + cached × cached rate. Kāore te tatau cached e nui ake i te tatau input.

Model Input / 1M Input cached / 1M
DeepSeek-V4-Pro-0813-3BIT-REAP $1.95 $0.488
GLM-5.2-3BIT-REAP $0.73 $0.183
Kimi-K3-3BIT-REAP $3.83 $0.958
Nemotron3Ultra-3BIT-REAP $0.75 $0.188
MiniMax-M3-3BIT-REAP $0.50 $0.125
DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP $0.50 $0.125
Kimi-K2.6-W4A16-AUTOROUND-REAP $0.78 $0.195
Laguna-S-2.1-W4A16-AUTOROUND-REAP $0.50 $0.125
inkling-W4A16-AUTOROUND-REAP $1.42 $0.355
MiMo-V2.5-Pro-W8A16 $0.50 $0.125
MiMo-V2.5-W8A16 $0.50 $0.125
Hy3-W8A16 $0.50 $0.125

Ka rārangi te usage log i te input cached o ia karanga. Kua uru kē ki ōna tokens utua me tōna utu te cached rate. Keys & usage

Ngā wāhi whakamahi

Endpoint Input cached Whakaaro
/v1/chat/completions usage.prompt_tokens_details.cached_tokens — wāhanga o te prompt_tokens usage.completion_tokens_details.reasoning_tokens — wāhanga o te completion_tokens
/v1/responses usage.input_tokens_details.cached_tokens — wāhanga o te input_tokens usage.output_tokens_details.reasoning_tokens — wāhanga o te output_tokens
/v1/messages usage.cache_read_input_tokens — ripana hei wāhi ake: ko te input_tokens te wāhanga ehara i te cached; ko te cache_creation_input_tokens he 0 i ngā wā katoa ka tatau te thinking i roto i ngā output_tokens
{
  "usage": {
    "prompt_tokens": 20000,
    "completion_tokens": 812,
    "total_tokens": 20812,
    "prompt_tokens_details": {
      "cached_tokens": 18000
    },
    "completion_tokens_details": {
      "reasoning_tokens": 604
    }
  }
}

Kei te whakautu stream ngā āpure ōrite i tōna usage whakamutunga. Kāore koe e tono:

Endpoint Te wāhi e tae mai ai te usage
/v1/chat/completions Te usage i te wāhanga whakamutunga i mua i te data: [DONE]. Ka tukuna i ia stream.
/v1/responses Te response.usage o te tuhinga response.completed.
/v1/messages Te usage o te tuhinga message_delta. Kei te usage o te message_start ngā kore.

Kia whakahau ake ngā cache hits

  • Kia mau te system prompt me ngā tohutohu tool kia tūmau i waenga i ngā calls. Whakawhitingia ngā wātau pēnei i ngā timestamps ki te mutunga o te karere hou, ehara i te system prompt.
  • Tāpiringia anake ki te hitori. Ki te whakatika, whakatuporo, neke rānei i ngā huringa mua, ka huri te prefix, ā, ka ututia katoa te wāhanga whai ake hei input noa.
  • Kaua e huri te huinga o ngā tools, ngā karere, neke rānei i ngā content blocks i waenga i ngā calls, ā, whakatikahia te JSON (tool schemas, tool arguments, me ngā results) i te wāhi kotahi i ngā wā katoa.
  • Noho ki tētahi model id mō tētahi kōrero, ā, tukuna te karanga whai muri i muri tata mai i te mea o mua.

Ka pupuri te API i te tīmatanga o tētahi kōrero kia pūmau i ēnei take:

  • Ka noho tētahi karere system, developer rānei i tukuna i muri i te kōrero ki tōna wāhi. Kāore e huri i te tīmatanga o te prompt, nō reira ka noho cached ngā tahuri i mua i a ia.
  • Ka whakatauritehia ngā tohenga o ngā tool call i ngā tahuri assistant o mua mā te uara. Kāore te raupapa key me te mokowā o taua JSON e pāngia.
  • Ka pānui ngā endpoint e toru i tētahi kōrero i te tikanga ōrite. Ka puritia e tētahi kōrero i haere tonu i tētahi atu endpoint tōna prefix tiritahi ina ōrite te ihirangi.

Ngā pāheke tono

Ko te prompt_cache_key (Chat Completions me ngā Responses) me te cache_control i ngā painga content blocks e uruhia ana, kia haere tonu ai te rongoa kōwae a te kiritaki. Ehara i te mea me whai: he automatika te caching, a mahi tonu tōna ahua ahua ahakao.

Āpure Tukuna ki He aha ia
prompt_cache_key /v1/chat/completions, /v1/responses He cache routing key o te API OpenAI.
cache_control /v1/messages He cache breakpoint i runga i tētahi poraka ihirangi, i tētahi poraka system, i tētahi karere rānei o te API Anthropic.
stream_options /v1/chat/completions Ka tono te include_usage i te API OpenAI mō te usage i runga i tētahi stream. Ki konei ka mutu ia stream me te usage.

Tupatu tokens

E rua ngā endpoint kore utu, POST /v1/tokenize me POST /v1/messages/count_tokens, ka tatau i ngā tokens o tētahi kuputuhi, o tētahi tono katoa rānei mō ngā model open-weight hosted i mua i tō tuku. He whārangi motuhake tō rāua: Te tatau token