Caching ea Prompt
KE AUTOMATICHosted open-weight models li cache-a prompt prefixes tse phetileng ka bo eona. Ha request e qala ka system prompt, tools le melaetsa e tšoanang le request ea morao-rao ho model e tšoanang, prefix eo e baloa ho tsoa ho cache ebile e betaloa ka 25% ea input price ea model. Ha ho na letho la ho enable, ebile cache writes ke mahala.
Mokhoa oa tšebetso
- Prefix, ka tatso — Prompt e baloa ka tatso: system prompt, tool definitions, ebe ke melaetsa. Cache e lekisa ho tloha qalong ea sequence eo ho fihla ho token ea pele e fapalaneng.
- Se queleloa hit — Request eo prompt ea eona e qalang ka content e tšoanang le request ea morao-rao — hangata ke turn ea feta ea conversation e tšoanang ka melaetsa e mecha e eketsitsoeng. Prefix e lekang ke cached input; tsohle ka morao oa eona ke regular input.
- Botlenya ba cache — Cache e boloka prompt ka li-block tsa tokens tse 1,568, kahoo prompt e khutšoanyane ho feta tokens tse ka bang 1,500 ha e cache-oe. Palo ea cached karabong ke palo ea hau ea input e atolositsoeng ka karolo ea cached ea prompt, e fokotsoeng ho tse fokolang. Ha e ame e le palo e atolositsoeng ea boholo ba block.
- Ntle le hit — Kopo eo qalo ea eona e seng ho cache e lefisoa ka input rate e tloaelehileng. Ha ho nako ea bophelo e phatlalalitsoeng bakeng sa li-prompt tse cached mme hit ha e tiisetsoe: bala
usageho bona seo kopo e se nkileng ho cache. - Ha ho switch — Kopo ha e kenelle, mme ha ho field e tima caching.
- Models lifelesso — Hosted open-weight id e mong’a. GET /v1/models e rapela capabilities.prompt_caching: true le pricing.cached_input_per_million_usd bakeng sa tsona. Shannon models li betala rate e le nngwe.
Bona cache hit karabong
Romela likopo tse peli tse qalang ka system prompt e telele e tšoanang mme u gatisetse usage ea e 'ngoe le e 'ngoe. Nomoro ea pele ke input ea kopo, ea bobeli ke karolo ea eona e baloang ho cache.
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")
handbook = open("handbook.txt").read() # a long text that stays the same
def ask(question):
response = client.chat.completions.create(
model="Kimi-K3-3BIT-REAP",
messages=[
{"role": "system", "content": handbook},
{"role": "user", "content": question},
],
)
usage = response.usage
print(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens)
ask("What is the refund policy?")
ask("Who approves travel?") # same start: read the second number import { readFileSync } from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({ apiKey: "YOUR_API_KEY", baseURL: "https://api.shannon-ai.com/v1" });
const handbook = readFileSync("handbook.txt", "utf8"); // a long text that stays the same
async function ask(question) {
const response = await client.chat.completions.create({
model: "Kimi-K3-3BIT-REAP",
messages: [
{ role: "system", content: handbook },
{ role: "user", content: question },
],
});
const usage = response.usage;
console.log(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens);
}
await ask("What is the refund policy?");
await ask("Who approves travel?"); // same start: read the second number # handbook.txt is a long text that stays the same. jq builds the JSON body from it
# and prints the usage object of the reply. Run it twice with different questions.
jq -Rs '{
model: "Kimi-K3-3BIT-REAP",
messages: [
{role: "system", content: .},
{role: "user", content: "What is the refund policy?"}
]
}' handbook.txt \
| curl -s https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d @- \
| jq .usage Theko
Cached input tokens li betaloa ka 25% ea input rate ea model, e kokotloang ho $0.001 per 1M. Ho ngola ho cache ha ho na lebeletso ea theko, ebile output e betaloa ka mokhoa oa tlhoahlo. Cached rate ea id ka’ng e teng tabuleng ea Models & pricing. Li-model le litheko
Input ea kopo e lefisoa e le (input − cached) × input rate + cached × cached rate. Palo ea cached ha e ka ke ea feta palo ea input.
| Model | Input / 1M | Input e cached / 1M |
|---|---|---|
DeepSeek-V4-Pro-0813-3BIT-REAP | $1.95 | $0.488 |
GLM-5.2-3BIT-REAP | $0.73 | $0.183 |
Kimi-K3-3BIT-REAP | $3.83 | $0.958 |
Nemotron3Ultra-3BIT-REAP | $0.75 | $0.188 |
MiniMax-M3-3BIT-REAP | $0.50 | $0.125 |
DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
Kimi-K2.6-W4A16-AUTOROUND-REAP | $0.78 | $0.195 |
Laguna-S-2.1-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
inkling-W4A16-AUTOROUND-REAP | $1.42 | $0.355 |
MiMo-V2.5-Pro-W8A16 | $0.50 | $0.125 |
MiMo-V2.5-W8A16 | $0.50 | $0.125 |
Hy3-W8A16 | $0.50 | $0.125 |
Usage log e thathamisa input e cached ea kopo e 'ngoe le e 'ngoe. Tokens tse lefisitsoeng le theko ea eona li se li kenyelletsa cached rate. Keys & usage
Fields tsa tšebeliso
| Endpoint | Input e cached | Reasoning |
|---|---|---|
/v1/chat/completions | usage.prompt_tokens_details.cached_tokens — karolo ea prompt_tokens | usage.completion_tokens_details.reasoning_tokens — karolo ea completion_tokens |
/v1/responses | usage.input_tokens_details.cached_tokens — karolo ea input_tokens | usage.output_tokens_details.reasoning_tokens — karolo ea output_tokens |
/v1/messages | usage.cache_read_input_tokens — e rapeloa ka setala: input_tokens ke karolo e sa cached-oeng; cache_creation_input_tokens ke 0 kamehla | thinking e baloa ho output_tokens |
{
"usage": {
"prompt_tokens": 20000,
"completion_tokens": 812,
"total_tokens": 20812,
"prompt_tokens_details": {
"cached_tokens": 18000
},
"completion_tokens_details": {
"reasoning_tokens": 604
}
}
} {
"usage": {
"input_tokens": 20000,
"input_tokens_details": {
"cached_tokens": 18000
},
"output_tokens": 812,
"output_tokens_details": {
"reasoning_tokens": 604
},
"total_tokens": 20812
}
} {
"usage": {
"input_tokens": 2000,
"cache_read_input_tokens": 18000,
"cache_creation_input_tokens": 0,
"output_tokens": 812
}
} Karabo ea stream e nka li-field tse tšoanang usage ea eona ea ho qetela. Ha u hloke ho e kopa:
| Endpoint | Moo usage e fihlang teng |
|---|---|
/v1/chat/completions | usage ho chunk ea ho qetela pele ho data: [DONE]. E romeloa ho stream e 'ngoe le e 'ngoe. |
/v1/responses | response.usage ea event ea response.completed. |
/v1/messages | usage ea event ea message_delta. usage ea message_start e na le li-zero. |
Ho fumana cache hits tse ngata
- Boloka system prompt le tool definitions li tšoana byte-ka-byte pakeng tsa likhoapolo. Beha values tsa call ka’ng joalo ka timestamps kapa request ids qetellong ea melaetsa e morao, eseng ho system prompt.
- Eketsa feela ho history. Ho fetola, ho fokotsa kapa ho summarize turns tsa pele ho fetola prefix, ebile tsohle ka morao oa pheto ea pele li betaloa e le regular input.
- u se fetole tatso ea tools, melaetsa kapa content blocks pakeng tsa likhoapolo, ebile serialise JSON (tool schemas, tool arguments le results) ka mokhoa o tšoanang kamehla.
- Lula ho id e le 'ngoe ea model moqoqong, mme u romele kopo e latelang kapele ka mor'a e fetileng.
API e boloka qalo ea moqoqo e tsitsitse maemong ana:
- Molaetsa oa
systemkapadevelopero romeloang hamorao moqoqong o lula sebakeng sa oona. Ha o fetole qalo ea prompt, kahoo liphetoho tse o fetileng li lula li cached. - Li-argument tsa likopo tsa tool liphetohong tse fetileng tsa assistant li bapisoa ka boleng. Tatellano ea li-key le sebaka sa JSON eo ha li bohlokoa.
- Li-endpoint tse tharo li bala moqoqo ka tsela e tšoanang. Moqoqo o tsoelang pele ho endpoint e 'ngoe o boloka prefix ea oona e arolelanoang ha litaba li tšoana.
Fields tsa kopo
prompt_cache_key (Chat Completions le Responses) le cache_control ho Message content blocks li ammehiloe, kahoo code ea client e seng elahlehile e sebetsa ntle le liphetwo. Ha ho na le eona ea hlokang: caching ke ea automatiki ’me e sebetsa ka tsela e tšoanang ntle le tsona.
| Field | E rometsoe ho | Ke eng |
|---|---|---|
prompt_cache_key | /v1/chat/completions, /v1/responses | Cache routing key ea OpenAI API. |
cache_control | /v1/messages | Cache breakpoint ho content block, block ea system kapa molaetsa oa Anthropic API. |
stream_options | /v1/chat/completions | include_usage e kopa OpenAI API usage ho stream. Mona stream e 'ngoe le e 'ngoe e fela ka usage. |
Ho bala tokens
Li-endpoint tse peli tsa mahala, POST /v1/tokenize le POST /v1/messages/count_tokens, li bala tokens tsa mongolo kapa tsa kopo eohle bakeng sa li-model tse hosted tsa open-weight pele u e romela. Li na le leqephe la tsona: Ho bala tokens