የፕሮምፕት ካሽንግ
CΑUTOMATICየhosted open-weight ሞዴሎች ተደጋጋሚ የprompt prefixesን በራስ-ሰር ያከማቻሉ። አንድ ጥያቄ በቅርብ ጊዜ በዚያው ሞዴል ላይ ጥቅም ላይ ከዋለው ጥያቄ ጋር ተመሳሳይ የsystem prompt፣ tools እና የቀድሞ መልዕክቶች ካሉት፣ ያ የጋራ prefix ከcache ይነበባል እና በሞዴሉ የinput ዋጋ 25% ይታሰባል። የሚያነቃው ምንም ነገር የለም፣ cache writes ደግሞ በነጻ ነው።
እንዴት እንደሚሰራ
- Prefix, በቅደም ተከተል — Prompt በቅደም ተከተል ይነበባል፡ system prompt፣ tool definitions፣ ከዚያም መልዕክቶች። Cache የሚሰላው ከዚህ ቅደም ተከተል መጀመሪያ ጀምሮ የመጀመሪያው የተለየ ቶክን እስከሚገኝ ድረስ ነው።
- ምን እንደ hit ይቆጠራል — የአንድ ጥያቄ prompt በቅርብ ጊዜ በነበረ ጥያቄ ይዘት የሚጀምር ከሆነ — አብዛኛውን ጊዜ አዳዲስ መልዕክቶች የተጨመሩበት የዚያው የconversation ቀድሞ የነበረው ዙር። የሚዛመደው prefix cached input ነው፤ ከዚያ በኋላ ያለው ሁሉ regular input ነው።
- ዝርዝርነት (Granularity) — Cache prompt በ1,568 tokens blocks ይይዛል፣ ስለዚህ ከ1,500 tokens ገደማ የሚያንስ prompt cache አይደረግም። በምላሽ ውስጥ ያለው cached ብዛት የinput ብዛትዎ ሲባዛ በprompt cached ድርሻ ነው፣ ወደ ታች ተጠጋግቶ። የblock መጠን ብዜት መሆን አይጠበቅበትም።
- ያለ hit — መጀመሪያው በcache ውስጥ የሌለ ጥያቄ በተለመደው የinput ዋጋ ይሰላል። ለcached prompts የቆይታ ጊዜ አልታተመም እና hit ዋስትና የለውም፦ ጥያቄ ከcache ምን እንደወሰደ ለማየት
usageን ያንብቡ። - መቀየሪያ የለም — ጥያቄ አይመርጥም፣ እና cachingን የሚያጠፋ መስክ የለም።
- የትኞቹ ሞዴሎች — ሁሉም hosted open-weight idዎች። GET /v1/models capabilities.prompt_caching: true እና pricing.cached_input_per_million_usd ያሳውቃሉ። የShannon ሞዴሎች አንድ ወጥ ተመንን ይጠቀማሉ።
በምላሽ ውስጥ cache hit ይመልከቱ
በተመሳሳይ ረጅም system prompt የሚጀምሩ ሁለት ጥያቄዎችን ይላኩ እና የእያንዳንዱን usage ያትሙ። የመጀመሪያው ቁጥር የጥያቄው input ነው፣ ሁለተኛው ከcache የተነበበው ክፍሉ ነው።
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")
handbook = open("handbook.txt").read() # a long text that stays the same
def ask(question):
response = client.chat.completions.create(
model="Kimi-K3-3BIT-REAP",
messages=[
{"role": "system", "content": handbook},
{"role": "user", "content": question},
],
)
usage = response.usage
print(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens)
ask("What is the refund policy?")
ask("Who approves travel?") # same start: read the second number import { readFileSync } from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({ apiKey: "YOUR_API_KEY", baseURL: "https://api.shannon-ai.com/v1" });
const handbook = readFileSync("handbook.txt", "utf8"); // a long text that stays the same
async function ask(question) {
const response = await client.chat.completions.create({
model: "Kimi-K3-3BIT-REAP",
messages: [
{ role: "system", content: handbook },
{ role: "user", content: question },
],
});
const usage = response.usage;
console.log(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens);
}
await ask("What is the refund policy?");
await ask("Who approves travel?"); // same start: read the second number # handbook.txt is a long text that stays the same. jq builds the JSON body from it
# and prints the usage object of the reply. Run it twice with different questions.
jq -Rs '{
model: "Kimi-K3-3BIT-REAP",
messages: [
{role: "system", content: .},
{role: "user", content: "What is the refund policy?"}
]
}' handbook.txt \
| curl -s https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d @- \
| jq .usage ዋጋ
Cached input tokens በሞዴሉ የinput ተመን 25% ይታሰባሉ፣ ይህም በ 1M ቶክን $0.001 ተደርጎ ይበረራል። ወደ cache መጻፍ ምንም ተጨማሪ ወጪ የለውም፣ output ደግሞ እንደተለመደው ይታሰባል። የእያንዳንዱ id cached rate በModels & pricing ሰንጠረዥ ውስጥ ይገኛል። ሞዴሎችና ዋጋ
የጥሪ input የሚሰላው እንደ (input − cached) × የinput ዋጋ + cached × የcached ዋጋ ነው። የcached ብዛት ከinput ብዛት በፍጹም አይበልጥም።
| ሞዴል | Input / 1M | የcached input / 1M |
|---|---|---|
DeepSeek-V4-Pro-0813-3BIT-REAP | $1.95 | $0.488 |
GLM-5.2-3BIT-REAP | $0.73 | $0.183 |
Kimi-K3-3BIT-REAP | $3.83 | $0.958 |
Nemotron3Ultra-3BIT-REAP | $0.75 | $0.188 |
MiniMax-M3-3BIT-REAP | $0.50 | $0.125 |
DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
Kimi-K2.6-W4A16-AUTOROUND-REAP | $0.78 | $0.195 |
Laguna-S-2.1-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
inkling-W4A16-AUTOROUND-REAP | $1.42 | $0.355 |
MiMo-V2.5-Pro-W8A16 | $0.50 | $0.125 |
MiMo-V2.5-W8A16 | $0.50 | $0.125 |
Hy3-W8A16 | $0.50 | $0.125 |
የusage ምዝግብ የእያንዳንዱን ጥሪ cached input ይዘረዝራል። የተሰላው tokens እና ወጪው የcached ዋጋን አስቀድሞ ያካትታሉ። Keys & usage
የአጠቃቀም መስኮች
| ኤንድ-ፖይንት (Endpoint) | የተ-ካሽ input | ምክንያታዊነት (Reasoning) |
|---|---|---|
/v1/chat/completions | usage.prompt_tokens_details.cached_tokens — የ prompt_tokens አካል | usage.completion_tokens_details.reasoning_tokens — የ completion_tokens አካል |
/v1/responses | usage.input_tokens_details.cached_tokens — የ input_tokens አካል | usage.output_tokens_details.reasoning_tokens — የ output_tokens አካል |
/v1/messages | usage.cache_read_input_tokens — በተናጠል ሪፖርት የሚደረግ፡ input_tokens ያልተከማቸዉ ክፍል ነው፤ cache_creation_input_tokens ሁልጊዜ 0 ነው | thinking በ output_tokens ውስጥ ይቆጠራል |
{
"usage": {
"prompt_tokens": 20000,
"completion_tokens": 812,
"total_tokens": 20812,
"prompt_tokens_details": {
"cached_tokens": 18000
},
"completion_tokens_details": {
"reasoning_tokens": 604
}
}
} {
"usage": {
"input_tokens": 20000,
"input_tokens_details": {
"cached_tokens": 18000
},
"output_tokens": 812,
"output_tokens_details": {
"reasoning_tokens": 604
},
"total_tokens": 20812
}
} {
"usage": {
"input_tokens": 2000,
"cache_read_input_tokens": 18000,
"cache_creation_input_tokens": 0,
"output_tokens": 812
}
} የstream ምላሽ በመጨረሻ usageው ውስጥ ተመሳሳይ መስኮችን ይይዛል። መጠየቅ አያስፈልግዎትም፦
| ኤንድ-ፖይንት (Endpoint) | usage የሚደርስበት |
|---|---|
/v1/chat/completions | ከdata: [DONE] በፊት ባለው የመጨረሻ chunk ላይ usage። በእያንዳንዱ stream ላይ ይላካል። |
/v1/responses | የresponse.completed event response.usage። |
/v1/messages | የmessage_delta event usage። የmessage_start usage ዜሮዎችን ይይዛል። |
የበለጠ cache hits ለማግኘት
- በጥያቄዎች መካከል የsystem prompt እና የtool definitions በባይት ደረጃ ተመሳሳይ እንዲሆኑ ያድርጉ። እንደ timestamps ወይም request ids ያሉ ተለዋዋጭ እሴቶችን በsystem prompt ውስጥ ሳይሆን በመጨረሻው መልዕክት ላይ ያድርጉ።
- በታሪኩ ላይ መጨመር ብቻ ይለማመዱ። የቀድሞ ዙሮችን ማስተካከል፣ ማሳጠር ወይም ማጠቃለል prefix-ን ይለውጠዋል፣ እና ከመጀመሪያው ለውጥ በኋላ ያለው ሁሉ እንደ regular input ይታሰባል።
- በጥያቄዎች መካከል የtools፣ የመልዕክቶች ወይም የcontent blocks ቅደም ተከተልን አይለውጡ፣ እና JSON (tool schemas, tool arguments እና results) ሁልጊዜ በተመሳሳይ መንገድ serialize ያድርጉ።
- ለአንድ ውይይት በአንድ የሞዴል id ላይ ይቆዩ፣ እና ተከታዩን ጥሪ ከቀዳሚው ብዙም ሳይቆይ ይላኩ።
API የውይይቱን መጀመሪያ በእነዚህ ሁኔታዎች ውስጥ ቋሚ ያደርጋል፦
- በውይይት ውስጥ በኋላ የተላከ
systemወይምdeveloperመልእክት በቦታው ይቆያል። የprompt መጀመሪያን አይለውጥም፣ ስለዚህ ከእሱ በፊት ያሉት ዙሮች cached ሆነው ይቆያሉ። - በቀድሞ የassistant ዙሮች ውስጥ የtool ጥሪዎች arguments በእሴት ይነጻጸራሉ። የዚያ JSON የkey ቅደም ተከተል እና ክፍተት ግድ የለውም።
- ሦስቱ endpoints ውይይትን በተመሳሳይ መንገድ ያነባሉ። በሌላ endpoint የቀጠለ ውይይት ይዘቱ ተመሳሳይ ሲሆን የጋራ prefixውን ይይዛል።
የጥያቄ ፊልዶች
prompt_cache_key (ለChat Completions እና Responses) እና በMessages content blocks ላይ cache_control ተቀባይ ናቸው፣ ስለዚህ ነባር የclient ኮዶች ሳይቀየሩ ይሰራሉ። ሁለቱም የግድ አስፈላጊ አይደሉም፡ caching አውቶማቲክ ነው እና ያለእነሱም በተመሳሳይ ይ-ሰራል።
| መስክ | የሚላክበት | ምንድን ነው |
|---|---|---|
prompt_cache_key | /v1/chat/completions, /v1/responses | የOpenAI API ጥያቄዎችን ወደ አንድ cache የሚመራ ቁልፍ። |
cache_control | /v1/messages | የAnthropic API የይዘት block፣ የsystem block ወይም የመልእክት ላይ cache breakpoint። |
stream_options | /v1/chat/completions | include_usage OpenAI APIን በstream ላይ usage ይጠይቃል። እዚህ እያንዳንዱ stream በusage ያበቃል። |
የtokens ቆጠራ
ሁለት ነጻ endpoints፣ POST /v1/tokenize እና POST /v1/messages/count_tokens፣ ለሆስት open-weight ሞዴሎች ከመላክዎ በፊት የጽሑፍ ወይም የሙሉ ጥያቄ tokens ይቆጥራሉ። የራሳቸው ገጽ አላቸው፦ የToken ቆጠራ