Tūhono caching
AUNOAKo ngā hosted open-weight models e cache-tia ana ngā prefix prompt whakaherekaia a-automa tika. Ki te tīmataria he request ki te system prompt, ngā tools, me ngā karere rite ki tētahi request hou i te model kotahi, ka anau taua prefix mai i te cache, a ka ututia i te 25% o te utu input a te model. Kāore he mea hei whakahou, ā, he kore utu te tuhi ki te cache.
Me pēhea tana mahi
- Prefix, i te huinga — Ka anau te prompt i te huinga: system prompt, ngā tohutohu tool, nā reira ngā karere. Ka rite te cache mai i te tīmatanga o taua huinga ki te token tuatahi e hē ana.
- He aha te 'hit' — He request e tīmata ana tōna prompt ki te ihiraga rite ki tētahi request hou — ā-tikanga, ko te huringa mua o te kōrero kotahi me ngā karere hou i tāpiringia. Ko te prefix rite tēnei he input cached; ko-kete katoa whai ake he input noa.
- Taumata taipitopito — Ka pupuri te cache i tētahi prompt i ngā poraka o te 1,568 tokens, nō reira kāore e cache te prompt iti iho i te 1,500 tokens tata. Ko te tatau cached i tētahi whakautu ko tō tatau input whakarea ki te wāhanga cached o te prompt, whakahekea. Kāore e kore he whakarea o te rahi poraka.
- Kāore he hit — Ka utua te tono kāore tōna tīmatanga i te cache i te input rate auau. Kāore he wā ora e whakaputaina ana mō ngā prompt cached, ā, kāore te hit e whakapaingia: pānuitia te
usagehei kite i tā te tono i tango mai i te cache. - Kāore he pana — Kāore te tono e whakaae, ā, kāore he āpure hei whakawetohia te caching.
- No aupapa models — Katoa ngā hosted open-weight id. Ko te GET /v1/models e ripana ana i te capabilities.prompt_caching: true me te pricing.cached_input_per_million_usd mō rātou. Ko ngā Shannon models he reiti kotahi anau.
Kite i tētahi cache hit i tētahi whakautu
Tukuna e rua ngā tono e tīmata ana me te system prompt roa ōrite ka taka i te usage o ia tētahi. Ko te tau tuatahi te input o te tono, ko te tuarua te wāhanga i pānuitia mai i te cache.
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")
handbook = open("handbook.txt").read() # a long text that stays the same
def ask(question):
response = client.chat.completions.create(
model="Kimi-K3-3BIT-REAP",
messages=[
{"role": "system", "content": handbook},
{"role": "user", "content": question},
],
)
usage = response.usage
print(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens)
ask("What is the refund policy?")
ask("Who approves travel?") # same start: read the second number import { readFileSync } from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({ apiKey: "YOUR_API_KEY", baseURL: "https://api.shannon-ai.com/v1" });
const handbook = readFileSync("handbook.txt", "utf8"); // a long text that stays the same
async function ask(question) {
const response = await client.chat.completions.create({
model: "Kimi-K3-3BIT-REAP",
messages: [
{ role: "system", content: handbook },
{ role: "user", content: question },
],
});
const usage = response.usage;
console.log(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens);
}
await ask("What is the refund policy?");
await ask("Who approves travel?"); // same start: read the second number # handbook.txt is a long text that stays the same. jq builds the JSON body from it
# and prints the usage object of the reply. Run it twice with different questions.
jq -Rs '{
model: "Kimi-K3-3BIT-REAP",
messages: [
{role: "system", content: .},
{role: "user", content: "What is the refund policy?"}
]
}' handbook.txt \
| curl -s https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d @- \
| jq .usage Utu
Ka utu ngā cached input tokens i te 25% o te reiti input a te model, ka whakatau ki te $0.001 per 1M. Kāore he utu tari a te tuhi ki te cache, ā, ka utu te output hei tikanga noa. Kei te tēpu Models & pricing te reiti cached mō ia id. Ngā model me ngā utu
Ka utua te input o tētahi karanga hei (input − cached) × input rate + cached × cached rate. Kāore te tatau cached e nui ake i te tatau input.
| Model | Input / 1M | Input cached / 1M |
|---|---|---|
DeepSeek-V4-Pro-0813-3BIT-REAP | $1.95 | $0.488 |
GLM-5.2-3BIT-REAP | $0.73 | $0.183 |
Kimi-K3-3BIT-REAP | $3.83 | $0.958 |
Nemotron3Ultra-3BIT-REAP | $0.75 | $0.188 |
MiniMax-M3-3BIT-REAP | $0.50 | $0.125 |
DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
Kimi-K2.6-W4A16-AUTOROUND-REAP | $0.78 | $0.195 |
Laguna-S-2.1-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
inkling-W4A16-AUTOROUND-REAP | $1.42 | $0.355 |
MiMo-V2.5-Pro-W8A16 | $0.50 | $0.125 |
MiMo-V2.5-W8A16 | $0.50 | $0.125 |
Hy3-W8A16 | $0.50 | $0.125 |
Ka rārangi te usage log i te input cached o ia karanga. Kua uru kē ki ōna tokens utua me tōna utu te cached rate. Keys & usage
Ngā wāhi whakamahi
| Endpoint | Input cached | Whakaaro |
|---|---|---|
/v1/chat/completions | usage.prompt_tokens_details.cached_tokens — wāhanga o te prompt_tokens | usage.completion_tokens_details.reasoning_tokens — wāhanga o te completion_tokens |
/v1/responses | usage.input_tokens_details.cached_tokens — wāhanga o te input_tokens | usage.output_tokens_details.reasoning_tokens — wāhanga o te output_tokens |
/v1/messages | usage.cache_read_input_tokens — ripana hei wāhi ake: ko te input_tokens te wāhanga ehara i te cached; ko te cache_creation_input_tokens he 0 i ngā wā katoa | ka tatau te thinking i roto i ngā output_tokens |
{
"usage": {
"prompt_tokens": 20000,
"completion_tokens": 812,
"total_tokens": 20812,
"prompt_tokens_details": {
"cached_tokens": 18000
},
"completion_tokens_details": {
"reasoning_tokens": 604
}
}
} {
"usage": {
"input_tokens": 20000,
"input_tokens_details": {
"cached_tokens": 18000
},
"output_tokens": 812,
"output_tokens_details": {
"reasoning_tokens": 604
},
"total_tokens": 20812
}
} {
"usage": {
"input_tokens": 2000,
"cache_read_input_tokens": 18000,
"cache_creation_input_tokens": 0,
"output_tokens": 812
}
} Kei te whakautu stream ngā āpure ōrite i tōna usage whakamutunga. Kāore koe e tono:
| Endpoint | Te wāhi e tae mai ai te usage |
|---|---|
/v1/chat/completions | Te usage i te wāhanga whakamutunga i mua i te data: [DONE]. Ka tukuna i ia stream. |
/v1/responses | Te response.usage o te tuhinga response.completed. |
/v1/messages | Te usage o te tuhinga message_delta. Kei te usage o te message_start ngā kore. |
Kia whakahau ake ngā cache hits
- Kia mau te system prompt me ngā tohutohu tool kia tūmau i waenga i ngā calls. Whakawhitingia ngā wātau pēnei i ngā timestamps ki te mutunga o te karere hou, ehara i te system prompt.
- Tāpiringia anake ki te hitori. Ki te whakatika, whakatuporo, neke rānei i ngā huringa mua, ka huri te prefix, ā, ka ututia katoa te wāhanga whai ake hei input noa.
- Kaua e huri te huinga o ngā tools, ngā karere, neke rānei i ngā content blocks i waenga i ngā calls, ā, whakatikahia te JSON (tool schemas, tool arguments, me ngā results) i te wāhi kotahi i ngā wā katoa.
- Noho ki tētahi model id mō tētahi kōrero, ā, tukuna te karanga whai muri i muri tata mai i te mea o mua.
Ka pupuri te API i te tīmatanga o tētahi kōrero kia pūmau i ēnei take:
- Ka noho tētahi karere
system,developerrānei i tukuna i muri i te kōrero ki tōna wāhi. Kāore e huri i te tīmatanga o te prompt, nō reira ka noho cached ngā tahuri i mua i a ia. - Ka whakatauritehia ngā tohenga o ngā tool call i ngā tahuri assistant o mua mā te uara. Kāore te raupapa key me te mokowā o taua JSON e pāngia.
- Ka pānui ngā endpoint e toru i tētahi kōrero i te tikanga ōrite. Ka puritia e tētahi kōrero i haere tonu i tētahi atu endpoint tōna prefix tiritahi ina ōrite te ihirangi.
Ngā pāheke tono
Ko te prompt_cache_key (Chat Completions me ngā Responses) me te cache_control i ngā painga content blocks e uruhia ana, kia haere tonu ai te rongoa kōwae a te kiritaki. Ehara i te mea me whai: he automatika te caching, a mahi tonu tōna ahua ahua ahakao.
| Āpure | Tukuna ki | He aha ia |
|---|---|---|
prompt_cache_key | /v1/chat/completions, /v1/responses | He cache routing key o te API OpenAI. |
cache_control | /v1/messages | He cache breakpoint i runga i tētahi poraka ihirangi, i tētahi poraka system, i tētahi karere rānei o te API Anthropic. |
stream_options | /v1/chat/completions | Ka tono te include_usage i te API OpenAI mō te usage i runga i tētahi stream. Ki konei ka mutu ia stream me te usage. |
Tupatu tokens
E rua ngā endpoint kore utu, POST /v1/tokenize me POST /v1/messages/count_tokens, ka tatau i ngā tokens o tētahi kuputuhi, o tētahi tono katoa rānei mō ngā model open-weight hosted i mua i tō tuku. He whārangi motuhake tō rāua: Te tatau token