Prompt caching
အလိုအလျောက်Hosted open-weight models များသည် ထပ်တလဲလဲဖြစ်သော prompt prefixes များကို အလိုအလျောက် cache လုပ်ပါသည်။ Request တစ်ခုသည် ယခင် request တစ်ခုနှင့် တူညီသော system prompt, tools နှင့် message များဖြင့် စတင်ပါက၊ ၎င်း shared prefix ကို cache မှ ဖတ်ရှုပြီး model ၏ input ဈေးနှုန်း ၂၅% ဖြင့်သာ ပေးချေရပါသည်။ သီးခြား enable လုပ်ရန် မလိုအပ်ဘဲ cache write လုပ်ခြင်းသည် အခမဲ့ ဖြစ်ပါသည်။
အလုပ်လုပ်ပုံ
- Prefix, အစဉ်လိုက် — Prompt ကို အစီအစဉ်အတိုင်း ဖတ်ရှုသည်- system prompt, tool definitions နှင့် messages တို့ ဖြစ်ပါသည်။ Cache သည် ထိုအစီအစဉ်၏ အစမှ ပထမဆုံး ကွဲပြားသွားသော token အထိ ကိုက်ညီမှုရှိရပါမည်။
- Hit အဖြစ် သတ်မှတ်ခြင်း — Request တစ်ခု၏ prompt သည် မကြာသေးမီက ပေးပို့ခဲ့သော request နှင့် အကြောင်းအရာ တူညီနေခြင်း ဖြစ်သည် — ပုံမှန်အားဖြင့် message အသစ်များ ဖြည့်စွက်ထားသော တူညီသည့် conversation ၏ ယခင်အလှည့် ဖြစ်ပါသည်။ ကိုက်ညီသော prefix သည် cached input ဖြစ်ပြီး ၎င်းနောက်ရှိ အရာအားလုံးသည် regular input ဖြစ်ပါသည်။
- အသေးစိတ်အဆင့် (Granularity) — Cache သည် prompt ကို 1,568 token ပါသော block များဖြင့် သိမ်းထားသဖြင့် token 1,500 ခန့်ထက်တိုသော prompt ကို cache မလုပ်ပါ။ Reply ထဲရှိ cached အရေအတွက်သည် သင့် input အရေအတွက်ကို prompt ၏ cached အချိုးနှင့် မြှောက်ပြီး အောက်သို့ ဖြတ်ထားသော တန်ဖိုးဖြစ်သည်။ Block အရွယ်အစား၏ ဆတိုးဖြစ်ရန် မလိုအပ်ပါ။
- Hit မရှိလျှင် — အစပိုင်းသည် cache ထဲတွင် မရှိသော request ကို ပုံမှန် input နှုန်းဖြင့် ကောက်ခံသည်။ Cache လုပ်ထားသော prompt များအတွက် သက်တမ်းကို မထုတ်ပြန်ထားသလို hit ဖြစ်မည်ဟု အာမမခံပါ- request တစ်ခုက cache မှ ယူခဲ့သည်ကို
usageတွင် ဖတ်ပါ။ - ခလုတ် မရှိပါ — Request တစ်ခုက ရွေးချယ်ပေးရန် မလိုသလို caching ကို ပိတ်သော field လည်း မရှိပါ။
- မည်သည့် model များ — Hosted open-weight id အားလုံးတွင် အသုံးပြုနိုင်ပါသည်။ GET /v1/models တွင် capabilities.prompt_caching: true နှင့် pricing.cached_input_per_million_usd တို့ကို ဖော်ပြထားပါသည်။ Shannon models များသည် flat rate တစ်ခုတည်းဖြင့် တွက်ချက်ပါသည်။
Reply တွင် cache hit ကို ကြည့်ရန်
တူညီသော ရှည်လျားသည့် system prompt ဖြင့်စသော request နှစ်ခုကို ပို့ပြီး တစ်ခုစီ၏ usage ကို print ထုတ်ပါ။ ပထမနံပါတ်သည် request ၏ input ဖြစ်ပြီး ဒုတိယနံပါတ်သည် ၎င်းထဲမှ cache မှ ဖတ်ထားသော အပိုင်း ဖြစ်သည်။
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")
handbook = open("handbook.txt").read() # a long text that stays the same
def ask(question):
response = client.chat.completions.create(
model="Kimi-K3-3BIT-REAP",
messages=[
{"role": "system", "content": handbook},
{"role": "user", "content": question},
],
)
usage = response.usage
print(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens)
ask("What is the refund policy?")
ask("Who approves travel?") # same start: read the second number import { readFileSync } from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({ apiKey: "YOUR_API_KEY", baseURL: "https://api.shannon-ai.com/v1" });
const handbook = readFileSync("handbook.txt", "utf8"); // a long text that stays the same
async function ask(question) {
const response = await client.chat.completions.create({
model: "Kimi-K3-3BIT-REAP",
messages: [
{ role: "system", content: handbook },
{ role: "user", content: question },
],
});
const usage = response.usage;
console.log(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens);
}
await ask("What is the refund policy?");
await ask("Who approves travel?"); // same start: read the second number # handbook.txt is a long text that stays the same. jq builds the JSON body from it
# and prints the usage object of the reply. Run it twice with different questions.
jq -Rs '{
model: "Kimi-K3-3BIT-REAP",
messages: [
{role: "system", content: .},
{role: "user", content: "What is the refund policy?"}
]
}' handbook.txt \
| curl -s https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d @- \
| jq .usage ဈေးနှုန်းသတ်မှတ်ချက်
Cached input tokens များကို model ၏ input rate ၂၅% နှုန်းဖြင့် တွက်ချက်ပြီး 1M လျှင် $0.001 အထိ ညှိထားပါသည်။ Cache သို့ ရေးသားခြင်းအတွက် အပိုကုန်ကျစရိတ်မရှိဘဲ output ကို ပုံမှန်အတိုင်း ပေးချေရပါသည်။ id တစ်ခုချင်းစီ၏ cached rate ကို Models & pricing table တွင် ကြည့်ရှုနိုင်ပါသည်။ Model များနှင့် ဈေးနှုန်း
Call တစ်ခု၏ input ကို (input − cached) × input နှုန်း + cached × cached နှုန်း ဖြင့် ကောက်ခံသည်။ Cached အရေအတွက်သည် input အရေအတွက်ထက် ဘယ်တော့မှ မကြီးပါ။
| မော်ဒယ် | Input / 1M | Cache လုပ်ထားသော input / 1M |
|---|---|---|
DeepSeek-V4-Pro-0813-3BIT-REAP | $1.95 | $0.488 |
GLM-5.2-3BIT-REAP | $0.73 | $0.183 |
Kimi-K3-3BIT-REAP | $3.83 | $0.958 |
Nemotron3Ultra-3BIT-REAP | $0.75 | $0.188 |
MiniMax-M3-3BIT-REAP | $0.50 | $0.125 |
DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
Kimi-K2.6-W4A16-AUTOROUND-REAP | $0.78 | $0.195 |
Laguna-S-2.1-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
inkling-W4A16-AUTOROUND-REAP | $1.42 | $0.355 |
MiMo-V2.5-Pro-W8A16 | $0.50 | $0.125 |
MiMo-V2.5-W8A16 | $0.50 | $0.125 |
Hy3-W8A16 | $0.50 | $0.125 |
Usage log တွင် call တစ်ခုစီ၏ cache လုပ်ထားသော input ကို ဖော်ပြသည်။ ၎င်း၏ ကောက်ခံထားသော token များနှင့် ကုန်ကျစရိတ်တွင် cached နှုန်းကို ထည့်သွင်းပြီးဖြစ်သည်။ Keys & usage
အသုံးပြုမှုနယ်ပယ်များ
| Endpoint | Cache လုပ်ထားသော input | Reasoning |
|---|---|---|
/v1/chat/completions | usage.prompt_tokens_details.cached_tokens — prompt_tokens ၏ အစိတ်အပိုင်း | usage.completion_tokens_details.reasoning_tokens — completion_tokens ၏ အစိတ်အပိုင်း |
/v1/responses | usage.input_tokens_details.cached_tokens — input_tokens ၏ အစိတ်အပိုင်း | usage.output_tokens_details.reasoning_tokens — output_tokens ၏ အစိတ်အပိုင်း |
/v1/messages | usage.cache_read_input_tokens — reported apart: input_tokens သည် uncached အပိုင်းဖြစ်ပြီး; cache_creation_input_tokens သည် အမြဲတမ်း 0 ဖြစ်သည် | thinking ကို output_tokens တွင် တွက်ချက်သည် |
{
"usage": {
"prompt_tokens": 20000,
"completion_tokens": 812,
"total_tokens": 20812,
"prompt_tokens_details": {
"cached_tokens": 18000
},
"completion_tokens_details": {
"reasoning_tokens": 604
}
}
} {
"usage": {
"input_tokens": 20000,
"input_tokens_details": {
"cached_tokens": 18000
},
"output_tokens": 812,
"output_tokens_details": {
"reasoning_tokens": 604
},
"total_tokens": 20812
}
} {
"usage": {
"input_tokens": 2000,
"cache_read_input_tokens": 18000,
"cache_creation_input_tokens": 0,
"output_tokens": 812
}
} Stream လုပ်ထားသော reply သည် ၎င်း၏ နောက်ဆုံး usage တွင် တူညီသော field များ ပါဝင်သည်။ တောင်းဆိုရန် မလိုပါ-
| Endpoint | Usage ရောက်လာသည့်နေရာ |
|---|---|
/v1/chat/completions | data: [DONE] မတိုင်မီ နောက်ဆုံး chunk ပေါ်ရှိ usage။ Stream တိုင်းတွင် ပို့သည်။ |
/v1/responses | response.completed event ၏ response.usage။ |
/v1/messages | message_delta event ၏ usage။ message_start ၏ usage တွင် သုညများသာ ပါသည်။ |
Cache hits ပိုမိုရရှိစေရန်
- Call များအကြား system prompt နှင့် tool definitions များကို byte-for-byte တူညီအောင် ထိန်းသိမ်းပါ။ timestamp သို့မဟုတ် request id ကဲ့သို့သော per-call တန်ဖိုးများကို system prompt တွင် မထားဘဲ နောက်ဆုံး message ၏ အဆုံးတွင် ထားပါ။
- History သို့ ဖြည့်စွက်ရုံသာ ပြုလုပ်ပါ။ ယခင်အလှည့်များကို ပြင်ဆင်ခြင်း၊ ဖြတ်တောက်ခြင်း သို့မဟုတ် အကျဉ်းချုပ်ခြင်းသည် prefix ကို ပြောင်းလဲစေပြီး ပထမဆုံး ပြောင်းလဲမှုနောက်ရှိ အရာအားလုံးကို regular input အဖြစ် ပေးချေရမည် ဖြစ်ပါသည်။
- Call များအကြား tools, messages သို့မဟုတ် content blocks များကို အစီအစဉ်မပြောင်းပါနှင့်၊ JSON (tool schemas, tool arguments နှင့် results) များကို အကြိမ်တိုင်း တူညီသောနည်းလမ်းဖြင့် serialise လုပ်ပါ။
- စကားပြောဆိုမှုတစ်ခုအတွက် model id တစ်ခုတည်းကို ဆက်သုံးပြီး နောက်ဆက်တွဲ call ကို ယခင် call ပြီးနောက် မကြာမီ ပို့ပါ။
API သည် ဤအခြေအနေများတွင် စကားပြောဆိုမှု၏ အစကို တည်ငြိမ်စွာ ထိန်းထားသည်-
- စကားပြောဆိုမှု နောက်ပိုင်းတွင် ပို့သော
systemသို့မဟုတ်developermessage သည် ၎င်း၏နေရာတွင်ပင် ရှိနေသည်။ ၎င်းသည် prompt ၏ အစကို မပြောင်းလဲသဖြင့် ၎င်းမတိုင်မီ turn များသည် cache ထဲတွင် ဆက်ရှိသည်။ - အစောပိုင်း assistant turn များရှိ tool call များ၏ argument များကို တန်ဖိုးဖြင့် နှိုင်းယှဉ်သည်။ ထို JSON ၏ key အစီအစဉ်နှင့် အကွာအဝေးသည် အရေးမကြီးပါ။
- Endpoint သုံးခုလုံးသည် စကားပြောဆိုမှုကို တူညီစွာ ဖတ်သည်။ အခြား endpoint တွင် ဆက်လက်ပြုလုပ်သော စကားပြောဆိုမှုသည် အကြောင်းအရာတူလျှင် မျှဝေထားသော prefix ကို ဆက်ထိန်းထားသည်။
Request fields များ
prompt_cache_key (Chat Completions နှင့် Responses) နှင့် Messages content blocks များရှိ cache_control တို့ကို လက်ခံသည်၊ ထို့ကြောင့် လက်ရှိ client code များကို ပြောင်းလဲရန်မလိုဘဲ အသုံးပြုနိုင်ပါသည်။ ၎င်းတို့အားလုံး မလိုအပ်ပါ— caching သည် အလိုအလျောက် အလုပ်လုပ်ပြီး ၎င်းတို့မပါဘဲလည်း ပုံမှန်အတိုင်း အလုပ်လုပ်ပါသည်။
| Field | ပို့သည့်နေရာ | အဓိပ္ပာယ် |
|---|---|---|
prompt_cache_key | /v1/chat/completions, /v1/responses | OpenAI API ၏ cache routing key ဖြစ်ပြီး cache ကို လမ်းညွှန်ရန် အသုံးပြုသည်။ |
cache_control | /v1/messages | Anthropic API တွင် content block၊ system block သို့မဟုတ် message တစ်ခုပေါ်၌ ထားသော cache breakpoint (cache ကို ဖြတ်ရန် အမှတ်အသား) ဖြစ်သည်။ |
stream_options | /v1/chat/completions | include_usage သည် OpenAI API တွင် stream ပေါ်၌ usage ကို တောင်းသည်။ ဤနေရာတွင် stream တိုင်းသည် usage ဖြင့် အဆုံးသတ်သည်။ |
Tokens များကို ရေတွက်ခြင်း
အခမဲ့ endpoint နှစ်ခုဖြစ်သော POST /v1/tokenize နှင့် POST /v1/messages/count_tokens တို့က မပို့မီ host လုပ်ထားသော open-weight မော်ဒယ်များအတွက် စာသားတစ်ခု သို့မဟုတ် request တစ်ခုလုံး၏ token များကို ရေတွက်ပေးသည်။ ၎င်းတို့တွင် သီးခြားစာမျက်နှာ ရှိသည်- Token ရေတွက်ခြင်း