การแคชพรอมต์
อัตโนมัติโมเดล hosted open-weight จะแคช prompt prefix ที่ซ้ำกันโดยอัตโนมัติ เมื่อคำขอเริ่มต้นด้วย system prompt, tools และข้อความก่อนหน้าเหมือนกับคำขอที่เพิ่งเกิดขึ้นในโมเดลเดียวกัน prefix ส่วนที่ใช้ร่วมกันนั้นจะถูกอ่านจากแคชและคิดค่าบริการเพียง 25% ของราคา input ของโมเดล ไม่ต้องตั้งค่าใดๆ และการเขียนลงแคชนั้นฟรี
หลักการทำงาน
- Prefix ตามลำดับ — Prompt จะถูกอ่านตามลำดับดังนี้: system prompt, tool definitions และตามด้วยข้อความ แคชจะจับคู่ตั้งแต่จุดเริ่มต้นของลำดับนั้นไปจนถึง token แรกที่แตกต่างกัน
- อะไรที่นับว่าเป็น hit — คำขอที่ prompt เริ่มต้นด้วยเนื้อหาเดียวกับคำขอที่เพิ่งเกิดขึ้น — โดยปกติคือการสนทนาในรอบก่อนหน้าที่มีการเพิ่มข้อความใหม่ต่อท้าย Prefix ที่ตรงกันคือ cached input; ส่วนที่เหลือหลังจากนั้นคือ regular input
- ความละเอียด (Granularity) — แคชเก็บ prompt เป็นบล็อกละ 1,568 tokens ดังนั้น prompt ที่สั้นกว่าประมาณ 1,500 tokens จะไม่ถูกแคช จำนวน cached ในคำตอบคือจำนวน input ของคุณคูณด้วยสัดส่วนของ prompt ที่แคชไว้ ปัดเศษลง ซึ่งไม่จำเป็นต้องเป็นพหุคูณของขนาดบล็อก
- เมื่อไม่เกิด hit — คำขอที่ส่วนต้นไม่อยู่ในแคชจะถูกคิดราคาตามอัตรา input ปกติ ไม่มีการประกาศอายุของ prompt ที่แคชไว้ และไม่รับประกันว่าจะเกิด hit: อ่าน
usageเพื่อดูว่าคำขอนั้นดึงอะไรมาจากแคช - ไม่มีสวิตช์ — คำขอไม่ต้องเลือกเข้าร่วม และไม่มีฟิลด์ใดปิดการแคชได้
- โมเดลที่รองรับ — ทุก hosted open-weight id โดย GET /v1/models จะรายงาน capabilities.prompt_caching: true และ pricing.cached_input_per_million_usd สำหรับโมเดลเหล่านี้ ส่วนโมเดล Shannon จะคิดอัตราเดียว
ดู cache hit ในคำตอบ
ส่งสองคำขอที่ขึ้นต้นด้วย system prompt ยาวเดียวกัน แล้วพิมพ์ usage ของแต่ละคำขอ ตัวเลขแรกคือ input ของคำขอ ตัวเลขที่สองคือส่วนที่อ่านจากแคช
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")
handbook = open("handbook.txt").read() # a long text that stays the same
def ask(question):
response = client.chat.completions.create(
model="Kimi-K3-3BIT-REAP",
messages=[
{"role": "system", "content": handbook},
{"role": "user", "content": question},
],
)
usage = response.usage
print(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens)
ask("What is the refund policy?")
ask("Who approves travel?") # same start: read the second number import { readFileSync } from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({ apiKey: "YOUR_API_KEY", baseURL: "https://api.shannon-ai.com/v1" });
const handbook = readFileSync("handbook.txt", "utf8"); // a long text that stays the same
async function ask(question) {
const response = await client.chat.completions.create({
model: "Kimi-K3-3BIT-REAP",
messages: [
{ role: "system", content: handbook },
{ role: "user", content: question },
],
});
const usage = response.usage;
console.log(usage.prompt_tokens, usage.prompt_tokens_details.cached_tokens);
}
await ask("What is the refund policy?");
await ask("Who approves travel?"); // same start: read the second number # handbook.txt is a long text that stays the same. jq builds the JSON body from it
# and prints the usage object of the reply. Run it twice with different questions.
jq -Rs '{
model: "Kimi-K3-3BIT-REAP",
messages: [
{role: "system", content: .},
{role: "user", content: "What is the refund policy?"}
]
}' handbook.txt \
| curl -s https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d @- \
| jq .usage ราคา
Cached input tokens คิดราคา 25% ของอัตรา input ของโมเดล ปัดเศษเป็น $0.001 ต่อ 1M การเขียนลงแคชไม่มีค่าใช้จ่ายเพิ่มเติม และ output คิดราคาตามปกติ อัตราแคชของแต่ละ id อยู่ในตาราง Models & pricing โมเดลและราคา
input ของการเรียกถูกคิดเป็น (input − cached) × อัตรา input + cached × อัตรา cached จำนวน cached ไม่เคยมากกว่าจำนวน input
| โมเดล | อินพุต / 1M | อินพุตที่แคช / 1M |
|---|---|---|
DeepSeek-V4-Pro-0813-3BIT-REAP | $1.95 | $0.488 |
GLM-5.2-3BIT-REAP | $0.73 | $0.183 |
Kimi-K3-3BIT-REAP | $3.83 | $0.958 |
Nemotron3Ultra-3BIT-REAP | $0.75 | $0.188 |
MiniMax-M3-3BIT-REAP | $0.50 | $0.125 |
DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
Kimi-K2.6-W4A16-AUTOROUND-REAP | $0.78 | $0.195 |
Laguna-S-2.1-W4A16-AUTOROUND-REAP | $0.50 | $0.125 |
inkling-W4A16-AUTOROUND-REAP | $1.42 | $0.355 |
MiMo-V2.5-Pro-W8A16 | $0.50 | $0.125 |
MiMo-V2.5-W8A16 | $0.50 | $0.125 |
Hy3-W8A16 | $0.50 | $0.125 |
บันทึกการใช้งานแสดง input ที่แคชของแต่ละการเรียก tokens ที่เรียกเก็บและค่าใช้จ่ายของบันทึกนั้นรวมอัตราของ cached ไว้แล้ว คีย์และการใช้งาน
ฟิลด์การใช้งาน
| เอนด์พอยต์ | อินพุตที่แคชไว้ | การให้เหตุผล |
|---|---|---|
/v1/chat/completions | usage.prompt_tokens_details.cached_tokens — ส่วนหนึ่งของ prompt_tokens | usage.completion_tokens_details.reasoning_tokens — ส่วนหนึ่งของ completion_tokens |
/v1/responses | usage.input_tokens_details.cached_tokens — ส่วนหนึ่งของ input_tokens | usage.output_tokens_details.reasoning_tokens — ส่วนหนึ่งของ output_tokens |
/v1/messages | usage.cache_read_input_tokens — รายงานแยกกัน: input_tokens คือส่วนที่ไม่ได้แคช; cache_creation_input_tokens จะเป็น 0 เสมอ | การคิด (thinking) ถูกนับรวมใน output_tokens |
{
"usage": {
"prompt_tokens": 20000,
"completion_tokens": 812,
"total_tokens": 20812,
"prompt_tokens_details": {
"cached_tokens": 18000
},
"completion_tokens_details": {
"reasoning_tokens": 604
}
}
} {
"usage": {
"input_tokens": 20000,
"input_tokens_details": {
"cached_tokens": 18000
},
"output_tokens": 812,
"output_tokens_details": {
"reasoning_tokens": 604
},
"total_tokens": 20812
}
} {
"usage": {
"input_tokens": 2000,
"cache_read_input_tokens": 18000,
"cache_creation_input_tokens": 0,
"output_tokens": 812
}
} คำตอบแบบสตรีมมีฟิลด์เดียวกันใน usage สุดท้าย ไม่ต้องร้องขอ:
| เอนด์พอยต์ | ที่ที่ usage มาถึง |
|---|---|
/v1/chat/completions | usage บน chunk สุดท้ายก่อน data: [DONE] ส่งในทุกสตรีม |
/v1/responses | response.usage ของอีเวนต์ response.completed |
/v1/messages | usage ของอีเวนต์ message_delta ส่วน usage ของ message_start เป็นศูนย์ทั้งหมด |
วิธีเพิ่มโอกาสเกิด cache hit
- รักษา system prompt และ tool definitions ให้คงเดิมแบบ byte-for-byte ในทุกการเรียกใช้ ให้ใส่ค่าที่เปลี่ยนตามการเรียก (เช่น timestamp หรือ request id) ไว้ที่ท้ายข้อความล่าสุด ไม่ใช่ใน system prompt
- ใช้วิธีการต่อท้าย (append) ประวัติการสนทนาเท่านั้น การแก้ไข ตัดทอน หรือสรุปเนื้อหาในรอบก่อนหน้าจะทำให้ prefix เปลี่ยน และทุกอย่างหลังจากจุดที่เปลี่ยนจะถูกคิดราคาเป็น regular input
- อย่าสลับลำดับของ tools, messages หรือ content blocks ระหว่างการเรียกใช้ และให้ serialize JSON (tool schemas, tool arguments และ results) ในรูปแบบเดียวกันทุกครั้ง
- ใช้ model id เดียวตลอดบทสนทนา และส่งการเรียกถัดไปในเวลาไม่นานหลังการเรียกก่อนหน้า
API ทำให้ส่วนต้นของบทสนทนาคงที่ในกรณีเหล่านี้:
- ข้อความ
systemหรือdeveloperที่ส่งภายหลังในบทสนทนาจะคงอยู่ที่ตำแหน่งเดิม ไม่เปลี่ยนส่วนต้นของ prompt ดังนั้นรอบก่อนหน้านั้นจึงยังคงถูกแคชอยู่ - อาร์กิวเมนต์ของการเรียก tool ในรอบ assistant ก่อนหน้าจะถูกเปรียบเทียบตามค่า ลำดับคีย์และช่องว่างของ JSON นั้นไม่มีผล
- ทั้งสาม endpoint อ่านบทสนทนาในลักษณะเดียวกัน บทสนทนาที่ต่อบน endpoint อื่นจะคง prefix ร่วมไว้ได้เมื่อเนื้อหาเหมือนกัน
ฟิลด์คำขอ (Request fields)
รองรับ prompt_cache_key (สำหรับ Chat Completions และ Responses) และ cache_control ใน content blocks ของ Messages ทำให้โค้ดเดิมทำงานได้โดยไม่ต้องแก้ไข ทั้งนี้ไม่จำเป็นต้องระบุทั้งสองอย่าง เนื่องจากระบบ caching ทำงานอัตโนมัติและให้ผลลัพธ์เหมือนกัน
| ฟิลด์ | ส่งไปที่ | คืออะไร |
|---|---|---|
prompt_cache_key | /v1/chat/completions, /v1/responses | คีย์กำหนดเส้นทางแคช (cache routing key) ของ OpenAI API |
cache_control | /v1/messages | จุดแบ่งแคช (cache breakpoint) บนบล็อกเนื้อหา บล็อก system หรือ message ของ Anthropic API |
stream_options | /v1/chat/completions | include_usage ใช้ขอ usage บนสตรีมจาก OpenAI API ที่นี่ทุกสตรีมจบด้วย usage อยู่แล้ว |
การนับ tokens
สอง endpoint ที่ไม่มีค่าใช้จ่าย คือ POST /v1/tokenize และ POST /v1/messages/count_tokens ใช้นับ tokens ของข้อความหรือของทั้งคำขอสำหรับโมเดล open-weight ที่โฮสต์ ก่อนที่คุณจะส่ง มีหน้าของตัวเอง: การนับ token