I-Prompt caching
IZIKHATHAI-hosted open-weight models i-cache i-prompt prefixes ephindaphindwayo ngokuzenzela. Xa i-request iqala nge-system prompt efanayo, i-tools kunye nemiyalezo efanayo ne-request yakuqhula kwi-model efanayo, loo prefix ifundwa kwi-cache ibhalwe ngama-25% wexabiso le-input le-model. Akukho nto kufuneka uyivule, kwaye ukubhala kwi-cache kungeza nantoni.
Isebenza njani
- Prefix, ngokulandelana — I-prompt ifundwa ngokulandelana: i-system prompt, i-tool definitions, emva koko imiyalezo. I-cache ihambelana ukusuka ekuqaleni kwelo sequence kuye kwi-token yokuqala efana nengafaniyo.
- Yintoni eyathathwa njenge-hit — I-request whose prompt iqala ngokuqukumbela okufanayo ne-request yakuqhula — ngokuhlala yinyanga yokuqala yengxoxo efanayo kunye nemiyalezo emitsha eyongeziweyo. I-prefix efanayo yi-cached input; yonke into emva koko yi-regular input.
- Ubuncwane — I-cache ibamba i-prompt kwiibloko ezingama-1,568 yee-token, ngoko i-prompt emfutshane kune-1,500 yee-token ngokuqikelelwa ayigcinwa. Inani le-cached kwimpendulo lelakho inani le-input lixhaswe yinxalenye egcinwe ye-prompt, lijikelwe ezantsi. Ayisoloko iphindaphindwa kobungakanani beebloko.
- Ngaphandle kwe-hit — Isicelo ekungekho ngqalelo yaso kwi-cache sihlawulwa kwireyithi eqhelekileyo ye-input. Akukho xesha lokuhlala lupapashiweyo lwee-prompt ezigcinwe kwaye i-hit ayiqinisekisiwe: funda i-
usageukubona oko isicelo sikuthathe kwi-cache. - Akukho sitshintshi — Isicelo asikhethi ukungena, kwaye akukho ntsimi icima i-caching.
- Zeziphi i-models — Yonke i-hosted open-weight id. I-GET /v1/models ibika i-capabilities.prompt_caching: true kunye ne-pricing.cached_input_per_million_usd kuzo. I-Shannon models ibhalwa ngemvuzo enye efanayo.
Bona i-cache hit kwimpendulo
Thumela izicelo ezimbini eziqala nge-system prompt ende efanayo uze uprinte ukusetyenziswa kwesinye ngasinye. Inani lokuqala yi-input yesicelo, lesibini yinxalenye yayo efundwe kwi-cache.
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")
handbook = open("handbook.txt").read() # a long text that stays the same
def ask(question):
response = client.chat.completions.create(
model="Kimi-K3-3BIT-REAP",
messages=[
{"role": "system", "content": handbook},
{"role": "user", "content": question},
],
)
usage = response.usage
print(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens)
ask("What is the refund policy?")
ask("Who approves travel?") # same start: read the second number import { readFileSync } from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({ apiKey: "YOUR_API_KEY", baseURL: "https://api.shannon-ai.com/v1" });
const handbook = readFileSync("handbook.txt", "utf8"); // a long text that stays the same
async function ask(question) {
const response = await client.chat.completions.create({
model: "Kimi-K3-3BIT-REAP",
messages: [
{ role: "system", content: handbook },
{ role: "user", content: question },
],
});
const usage = response.usage;
console.log(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens);
}
await ask("What is the refund policy?");
await ask("Who approves travel?"); // same start: read the second number # handbook.txt is a long text that stays the same. jq builds the JSON body from it
# and prints the usage object of the reply. Run it twice with different questions.
jq -Rs '{
model: "Kimi-K3-3BIT-REAP",
messages: [
{role: "system", content: .},
{role: "user", content: "What is the refund policy?"}
]
}' handbook.txt \
| curl -s https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d @- \
| jq .usage Amaxabiso
I-Cached input tokens zibhalwa ngama-25% wemvuzo ye-input le-model, i-rounded ibe ngu-$0.001 per 1M. Ukubhala kwi-cache akubizi nantoni eyongeziweyo, kwaye i-output ibhalwa ngok’Cala. Ixabiso le-cached le-id nganye likwita ye-Models & pricing. Iimodeli namaxabiso
I-input yobizo iyahlawulwa njenge (input − cached) × ireyithi ye-input + cached × ireyithi ye-cached. Inani le-cached alikho likhulu kuneli le-input.
| Imodeli | I-input / 1M | I-input egcinwe / 1M |
|---|---|---|
DeepSeek-V4-Pro-0813-3BIT-REAP | $1.95 | $0.488 |
GLM-5.2-3BIT-REAP | $0.73 | $0.183 |
Kimi-K3-3BIT-REAP | $3.83 | $0.958 |
Nemotron3Ultra-3BIT-REAP | $0.75 | $0.188 |
MiniMax-M3-3BIT-REAP | $0.50 | $0.125 |
DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
Kimi-K2.6-W4A16-AUTOROUND-REAP | $0.78 | $0.195 |
Laguna-S-2.1-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
inkling-W4A16-AUTOROUND-REAP | $1.42 | $0.355 |
MiMo-V2.5-Pro-W8A16 | $0.50 | $0.125 |
MiMo-V2.5-W8A16 | $0.50 | $0.125 |
Hy3-W8A16 | $0.50 | $0.125 |
Igosa lokusetyenziswa lidwelisa i-input egcinwe yobizo ngalunye. Ii-token zalo ezibhiliweyo nexabiso sele zibandakanya ireyithi ye-cached. Izitshixo nokusetyenziswa
Amacandelo okusetyenziswa
| Endpoint | I-input e-cached | Ukuqiqa |
|---|---|---|
/v1/chat/completions | usage.prompt_tokens_details.cached_tokens — inxalenye ye-prompt_tokens | usage.completion_tokens_details.reasoning_tokens — inxalenye ye-completion_tokens |
/v1/responses | usage.input_tokens_details.cached_tokens — inxalenye ye-input_tokens | usage.output_tokens_details.reasoning_tokens — inxalenye ye-output_tokens |
/v1/messages | usage.cache_read_input_tokens — ibikwa ngokwamaqela: i-input_tokens yinxalenye engacachiyelelo; i-cache_creation_input_tokens ihlala i-0 | ukucinga kubalwa kwi-output_tokens |
{
"usage": {
"prompt_tokens": 20000,
"completion_tokens": 812,
"total_tokens": 20812,
"prompt_tokens_details": {
"cached_tokens": 18000
},
"completion_tokens_details": {
"reasoning_tokens": 604
}
}
} {
"usage": {
"input_tokens": 20000,
"input_tokens_details": {
"cached_tokens": 18000
},
"output_tokens": 812,
"output_tokens_details": {
"reasoning_tokens": 604
},
"total_tokens": 20812
}
} {
"usage": {
"input_tokens": 2000,
"cache_read_input_tokens": 18000,
"cache_creation_input_tokens": 0,
"output_tokens": 812
}
} Impendulo ekwi-stream ithwala iintsimi ezifanayo kwi-usage yayo yokugqibela. Akufuneki uyicele:
| Endpoint | Apho ukusetyenziswa kufika khona |
|---|---|
/v1/chat/completions | I-usage kwi-chunk yokugqibela phambi kwe-data: [DONE]. Ithunyelwa kuyo yonke i-stream. |
/v1/responses | I-response.usage yesiganeko se-response.completed. |
/v1/messages | I-usage yesiganeko se-message_delta. I-usage ye-message_start ibamba ii-zero. |
Ukufumana i-cache hits ezinqabileyo
- Gcina i-system prompt kunye ne-tool definitions zingatshintshwa byte-for-byte phakathi kwe-calls. Faka i-values ezifana ne-timestamps okanye i-request ids ekugqibeleni komyalezo wokugqibela, hayi kwi-system prompt.
- Yongeza kuphela kwimbali (history). Ukulungisa, ukunciphisa okanye ukufupha imiyalezo yangaphambili kutshintsha i-prefix, kwaye yonke into emva koku kutshintshwa ibhalwa njenge-regular input.
- Ungatshintshi indlela ekuhlelwe ngayo i-tools, imiyalezo okanye i-content blocks phakathi kwe-calls, kwaye i-serialise i-JSON (tool schemas, tool arguments kunye ne-results) ngendlela efanayo rhoqo.
- Hlala kwi-id yemodeli enye kwincoko, uze uthumele ubizo olulandelayo kwakamsinya emva kwolwangaphambili.
I-API igcina ukuqala kwencoko kuzinzile kwezi meko:
- Umyalezo we-
systemokanye we-developerothunyelwe kamva kwincoko uhlala kwindawo yawo. Awutshintshi ukuqala kwe-prompt, ngoko iminyaka ephambi kwawo ihlala igcinwe. - Ii-argument zobizo lwee-tool kwiminyaka yangaphambili ye-assistant ziqhathaniswa ngexabiso. Ulandelelwano lwezitshixo nesithuba se-JSON leyo azibalulekanga.
- Ii-endpoint ezintathu zifunda incoko ngendlela enye. Incoko iqhubekile kwenye i-endpoint igcina i-prefix yayo ekwabelwana ngayo xa umxholo ufana.
Iifields zezicelo
i-prompt_cache_key (Chat Completions and Responses) kunye ne-cache_control kwiblocks zokuqukumbela kwe-Messages zamkelekile, ngoko ke ikhowudi yomsebenzi ekho ihamba ngaphandle kwenguqu.
| Intsimi | Ithunyelwe ku | Yintoni |
|---|---|---|
prompt_cache_key | /v1/chat/completions, /v1/responses | Isitshixo sokuhambisa i-cache se-OpenAI API. |
cache_control | /v1/messages | Indawo yokwahlula ye-cache kwi-block yomxholo, kwi-block ye-system okanye kumyalezo we-Anthropic API. |
stream_options | /v1/chat/completions | I-include_usage icela ukusetyenziswa kwi-OpenAI API kwi-stream. Apha yonke i-stream iphela ngokusetyenziswa. |
Ukubala ama-tokens
Ii-endpoint ezimbini zasimahla, i-POST /v1/tokenize ne-POST /v1/messages/count_tokens, zibala ii-token zombhalo okanye zesicelo sonke seemodeli ze-open-weight ezisingathiweyo phambi kokuba usithumele. Zinephepha lazo: Ukubala ii-token