I-Prompt caching
IZINZAKALOAma-hosted open-weight models acache-a ama-prompt prefixes aphindwayo ngokuzenzayo. Uma i-request iqala nge-system prompt efanayo, ama-tools nemilayezo efanayo ne-request eyakamuva kulo model, leyo prefix efanayo ifundwa ku-cache futhi ikhokhelwa ngama-25% we-input price yalo model. Akukho okufanele ukushinte, futhi ukubhala ku-cache kumahhala.
Isebenza kanjani
- Prefix, ngokulandelana — I-prompt ifundwa ngokulandelana: i-system prompt, ama-tool definitions, bese kulandela imilayezo. I-cache iyaklopha kusukela ekuqaleni kwalelo luhlelo kuze kufike ithokheni lokuqala elihlukile.
- Yini eyibalawe njenge-hit — I-request lapho i-prompt yayo iqala ngokuqukumbela okufanayo ne-request eyakamuva — imvamisa i-turn eyedlulile yengxoxo efanayo lapho kune-messages ezintsha ezengeziwe. I-prefix efanayo iy-cached input; yonke into emva kwalokhu iy-regular input.
- Ubuncane — I-cache igcina i-prompt ngama-block angama-token angu-1,568, ngakho i-prompt emfushane kuka-1,500 ama-token ayigcinwa. Inani eligcinwe empendulweni liyinani lakho lokufakwayo ligaxwe ngengxenye egcinwe ye-prompt, lijikelezwa ngokwehlisa. Akulona njalo isiphindaphindo sosayizi we-block.
- Ngaphandle kwe-hit — Isicelo esiqala sona singekho ku-cache sibiziswa ngezinga elivamile lokufakwayo. Akukho isikhathi sokuphila esishicilelwe kuma-prompt agcinwe futhi i-hit ayiqinisekiswa: funda i-
usageukubona okuthathwe isicelo ku-cache. - Asikho isishintshi — Isicelo asikhethi ukungena, futhi akukho field evala i-caching.
- Yimaphi ama-models — Wonke ama-hosted open-weight id. I-GET /v1/models ibika i-capabilities.prompt_caching: true ne-pricing.cached_input_per_million_usd ayo. Ama-Shannon models akhokhisa ngelizinga elilodwa elifanayo.
Bona i-cache hit empendulweni
Thumela izicelo ezimbili eziqala nge-system prompt efanayo emide bese uphrinta ukusetshenziswa kwesinye ngasinye. Inombolo yokuqala iwukufakwayo kwesicelo, eyesibili yingxenye yakho efundwe ku-cache.
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")
handbook = open("handbook.txt").read() # a long text that stays the same
def ask(question):
response = client.chat.completions.create(
model="Kimi-K3-3BIT-REAP",
messages=[
{"role": "system", "content": handbook},
{"role": "user", "content": question},
],
)
usage = response.usage
print(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens)
ask("What is the refund policy?")
ask("Who approves travel?") # same start: read the second number import { readFileSync } from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({ apiKey: "YOUR_API_KEY", baseURL: "https://api.shannon-ai.com/v1" });
const handbook = readFileSync("handbook.txt", "utf8"); // a long text that stays the same
async function ask(question) {
const response = await client.chat.completions.create({
model: "Kimi-K3-3BIT-REAP",
messages: [
{ role: "system", content: handbook },
{ role: "user", content: question },
],
});
const usage = response.usage;
console.log(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens);
}
await ask("What is the refund policy?");
await ask("Who approves travel?"); // same start: read the second number # handbook.txt is a long text that stays the same. jq builds the JSON body from it
# and prints the usage object of the reply. Run it twice with different questions.
jq -Rs '{
model: "Kimi-K3-3BIT-REAP",
messages: [
{role: "system", content: .},
{role: "user", content: "What is the refund policy?"}
]
}' handbook.txt \
| curl -s https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d @- \
| jq .usage Intengo
Ama-cached input tokens akhokhiselwa ngama-25% we-input rate yalo model, ajikelele ku-$0.001 per 1M. Ukubhala ku-cache akubizi lutho olwengeziwe, futhi i-output ikhokhelwa ngokujwayelekile. I-cached rate yalo mlawu isekhishini le-Models & pricing. Ama-model namanani
Okufakwayo kokubiza kubiziswa njenge (okufakwayo − okugcinwe) × izinga lokufakwayo + okugcinwe × izinga lokugcinwe. Inani eligcinwe alikaze libe likhulu kunenani lokufakwayo.
| I-model | Okufakwayo / 1M | Okufakwayo okugcinwe / 1M |
|---|---|---|
DeepSeek-V4-Pro-0813-3BIT-REAP | $1.95 | $0.488 |
GLM-5.2-3BIT-REAP | $0.73 | $0.183 |
Kimi-K3-3BIT-REAP | $3.83 | $0.958 |
Nemotron3Ultra-3BIT-REAP | $0.75 | $0.188 |
MiniMax-M3-3BIT-REAP | $0.50 | $0.125 |
DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
Kimi-K2.6-W4A16-AUTOROUND-REAP | $0.78 | $0.195 |
Laguna-S-2.1-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
inkling-W4A16-AUTOROUND-REAP | $1.42 | $0.355 |
MiMo-V2.5-Pro-W8A16 | $0.50 | $0.125 |
MiMo-V2.5-W8A16 | $0.50 | $0.125 |
Hy3-W8A16 | $0.50 | $0.125 |
Irekhodi lokusetshenziswa libala okufakwayo okugcinwe kwesicelo ngasinye. Ama-token abiziswayo nezindleko zalo sekufaka izinga lokugcinwe. Ama-key nokusetshenziswa
Ama-usage fields
| I-endpoint | Ingene ecashwe | Ukucabanga |
|---|---|---|
/v1/chat/completions | usage.prompt_tokens_details.cached_tokens — ingxenye ye-prompt_tokens | usage.completion_tokens_details.reasoning_tokens — ingxenye ye-completion_tokens |
/v1/responses | usage.input_tokens_details.cached_tokens — ingxenye ye-input_tokens | usage.output_tokens_details.reasoning_tokens — ingxenye ye-output_tokens |
/v1/messages | usage.cache_read_input_tokens — ibikwa ngokuhlukana: i-input_tokens iy-uncached part; i-cache_creation_input_tokens ihlezi i-0 | i-thinking ibalwa ku-output_tokens |
{
"usage": {
"prompt_tokens": 20000,
"completion_tokens": 812,
"total_tokens": 20812,
"prompt_tokens_details": {
"cached_tokens": 18000
},
"completion_tokens_details": {
"reasoning_tokens": 604
}
}
} {
"usage": {
"input_tokens": 20000,
"input_tokens_details": {
"cached_tokens": 18000
},
"output_tokens": 812,
"output_tokens_details": {
"reasoning_tokens": 604
},
"total_tokens": 20812
}
} {
"usage": {
"input_tokens": 2000,
"cache_read_input_tokens": 18000,
"cache_creation_input_tokens": 0,
"output_tokens": 812
}
} Impendulo esakazwayo iphethe ama-field afanayo ekusetshenzisweni kwayo kokugcina. Akudingeki uyicele:
| I-endpoint | Lapho ukusetshenziswa kufika khona |
|---|---|
/v1/chat/completions | I-usage ku-chunk yokugcina ngaphambi kwe-data: [DONE]. Ithunyelwa kuzo zonke izi-stream. |
/v1/responses | I-response.usage yomcimbi we-response.completed. |
/v1/messages | I-usage yomcimbi we-message_delta. I-usage ye-message_start iqukethe ama-zero. |
Ukuthola ama-cache hits amaningi
- Gcina i-system prompt nama-tool definitions engaguquki byte-for-byte phakathi kwe-calls. Faka izinto ezishintshayo njenge-timestamps noma ama-request ids ekugcineni kwemilayezo yakamuva, hhayi ku-system prompt.
- Yengeza kuphela emlandweni (history). Ukuhlela, ukufushane noma ukufingca ama-turns adlulile kushintsha i-prefix, futhi yonke into emva kokushintsha kukhokhelwa njenge-regular input.
- Ungaguquli ukuhleleka kwama-tools, imilayezo noma ama-content blocks phakathi kwe-calls, futhi hlela i-JSON (tool schemas, tool arguments ne-results) ngendlela efanayo ngaso sonke isikhathi.
- Hlala ku-id ye-model eyodwa engxoxweni, bese uthumela isicelo esilandelayo masinyane ngemva kwesandulele.
I-API igcina ukuqala kwengxoxo kuzinzile kule micimbi:
- Umlayezo we-
systemnoma we-developerothunyelwe kamuva engxoxweni uhlala endaweni yawo. Awushintshi ukuqala kwe-prompt, ngakho izinguquko ezandulela kuwo zihlala zigcinwe. - Izimpikiswano zokubiza amathuluzi ezinguqukweni zangaphambilini ze-assistant ziqhathaniswa ngenani. Ukuhleleka kwe-key nezikhala zalelo JSON akubalulekile.
- Ama-endpoint amathathu afunda ingxoxo ngendlela efanayo. Ingxoxo eqhutshekiswa kwenye i-endpoint igcina i-prefix yayo ehlanganyelwayo uma okuqukethwe kufana.
Ama-fields esicelo
I-prompt_cache_key (Chat Completions and Responses) kanye ne-cache_control kuma-Messages content blocks yamukeleka, ngakho i-client code ekhona isebenza ngaphandle kokushintshashutha. Akukho okudingekayo: i-caching yenzeka ngokuzenzela futhi isebenza ngendlela efanayo ngaphandle kwayo.
| I-field | Kuthunyelwa ku- | Yini |
|---|---|---|
prompt_cache_key | /v1/chat/completions, /v1/responses | I-cache routing key ye-OpenAI API. |
cache_control | /v1/messages | Indawo yokuhlukanisa ye-cache ku-block yokuqukethwe, ku-block ye-system noma emlayezweni we-Anthropic API. |
stream_options | /v1/chat/completions | I-include_usage icela ukusetshenziswa ku-OpenAI API ku-stream. Lapha zonke izi-stream ziphela ngokusetshenziswa. |
Ukubala ama-tokens
Ama-endpoint amabili ahlala mahhala, i-POST /v1/tokenize ne-POST /v1/messages/count_tokens, abala ama-token ombhalo noma wesicelo sonke kuma-model e-open-weight ahlinzekiwe ngaphambi kokuthumela. Zinekhasi lazo: Ukubala ama-token