Chat Completions
POST /v1/chat/completions รับบทสนทนาและคืนข้อความถัดไปของโมเดลในรูปแบบ OpenAI Chat Completions ใช้จาก OpenAI SDK ตัวใดก็ได้ หรือผ่าน HTTP ธรรมดา หน้านี้คือเอกสารอ้างอิงทีละฟิลด์
POST https://api.shannon-ai.com/v1/chat/completions
คำขอที่เล็กที่สุดคือ model id และข้อความของผู้ใช้หนึ่งข้อความ
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.shannon-ai.com/v1",
)
response = client.chat.completions.create(
model="shannon-3",
messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(response.choices[0].message.content) import OpenAI from "openai";
const client = new OpenAI({
apiKey: "YOUR_API_KEY",
baseURL: "https://api.shannon-ai.com/v1",
});
const response = await client.chat.completions.create({
model: "shannon-3",
messages: [{ role: "user", content: "Say hello in one sentence." }],
});
console.log(response.choices[0].message.content); curl https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "shannon-3",
"messages": [{"role": "user", "content": "Say hello in one sentence."}]
}' คำตอบคือออบเจกต์ JSON หนึ่งตัว:
{
"id": "chatcmpl-5f0c1e7a9b3d4c62a8e1f07d2b46c9a3",
"object": "chat.completion",
"created": 1791625200,
"model": "shannon-3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello, it is good to meet you.",
"reasoning_content": "The user wants a greeting in one sentence. Keep it short and friendly."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1184,
"completion_tokens": 46,
"total_tokens": 1230
}
} เฮดเดอร์
เฮดเดอร์ของคำขอ
| เฮดเดอร์ | ค่า | คำอธิบาย |
|---|---|---|
Authorization | Bearer YOUR_API_KEY | API key ของคุณ ใช้ x-api-key: YOUR_API_KEY แทนได้ในทุก endpoint |
Content-Type | application/json | จำเป็น ค่าอื่นใดจะคืน 415 |
x-request-id | ไม่บังคับ id ของคำขอที่คุณกำหนดเอง จะถูกส่งกลับมาในคำตอบโดยไม่เปลี่ยนแปลง |
เฮดเดอร์ของคำตอบ
| เฮดเดอร์ | คำอธิบาย |
|---|---|
x-request-id | อยู่ในทุกคำตอบ รวมถึงข้อผิดพลาดและสตรีม: ค่าที่คุณส่งมา หรืออักขระเลขฐานสิบหก 12 ตัวเมื่อคุณไม่ได้ส่ง ให้อ้างอิงค่านี้เมื่อรายงานปัญหา |
content-type | application/json หรือ text/event-stream เมื่อ stream เป็น true |
ฟิลด์คำขอ
ต้องระบุเฉพาะ messages คอลัมน์ ใช้โดย ระบุโมเดลที่ฟิลด์นั้นเปลี่ยนคำตอบได้ โมเดล open-weight ที่โฮสต์คือ id ทั้งสิบสองตัวในรายการโมเดล ส่วนตระกูล Shannon 3 คือ shannon-3, shannon-3-pro, shannon-3.1 และ shannon-3.1-pro โมเดลและราคา
| ฟิลด์ | ประเภท | ค่าเริ่มต้น | คำอธิบาย | ใช้โดย |
|---|---|---|---|---|
model | string | shannon-1.6-lite | โมเดลที่ตอบ: id จากรายการโมเดล ส่งมาทุกคำขอ การจับคู่ไม่สนใจตัวพิมพ์เล็กใหญ่ id ที่ไม่ได้เผยแพร่จะคืน 400 unknown model | ทุกโมเดล |
messages | array | จำเป็น บทสนทนา เรียงจากข้อความเก่าสุดก่อน ดู Messages ด้านล่าง | ทุกโมเดล | |
stream | boolean | false | true ส่งคำตอบเป็น server-sent events ระหว่างที่กำลังเขียน | ทุกโมเดล |
max_tokens | integer | 4096 | ขีดจำกัดสูงสุดของคำตอบ เป็นหน่วย tokens ค่าที่อยู่นอกช่วง 1 ถึง 65,536 จะถูกปรับให้อยู่ในช่วงนั้น ค่านี้ยังเป็นจำนวนที่กันไว้จากยอดคงเหลือของคุณในระหว่างที่คำขอทำงานอยู่ด้วย ดู ความยาวเอาต์พุต ด้านล่าง | โมเดล open-weight ที่โฮสต์, shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 |
max_completion_tokens | integer | เหมือนกับ max_tokens เมื่อส่งทั้งสองค่า จะใช้ max_tokens | โมเดล open-weight ที่โฮสต์, shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 | |
temperature | number | Sampling temperature บนโมเดล open-weight ที่โฮสต์ ค่าเริ่มต้นคือ 1 และค่าจะถูกจำกัดให้อยู่ระหว่าง 0 ถึง 2 | โมเดล open-weight ที่โฮสต์, shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 | |
top_p | number | 0.95 | Nucleus sampling ค่าจะถูกจำกัดให้อยู่ระหว่าง 0 ถึง 1 | โมเดล open-weight ที่โฮสต์ |
seed | integer | seed ของตัว sampler เป็นจำนวนเต็มใดก็ได้ หากไม่ระบุ seed จะคำนวณจากโมเดลและบทสนทนา ดังนั้นคำขอเดียวกันที่ส่งสองครั้งจะใช้ seed เดียวกัน | โมเดล open-weight ที่โฮสต์ | |
stop | string | array | สตริงหรืออาร์เรย์ของสตริง ใช้ได้สูงสุด 4 ตัว คำตอบจะจบก่อนตัวแรกที่ปรากฏ ข้อความหยุดนั้นเองจะไม่ถูกส่งกลับ | โมเดล open-weight ที่โฮสต์ | |
reasoning_effort | string | high | โมเดลให้เหตุผลมากเพียงใดก่อนตอบ: off, low, medium หรือ high โดย none และ minimal หมายถึง off, default หมายถึง medium และ max หมายถึง high ค่าอื่นใดจะคืน 400 | โมเดล open-weight ที่โฮสต์ |
reasoning | object | การตั้งค่าเดียวกันในรูปแบบออบเจกต์: {"effort": "low"} เมื่อส่งทั้งสองค่า จะใช้ reasoning_effort | โมเดล open-weight ที่โฮสต์ | |
tools | array | ฟังก์ชันที่โมเดลเรียกใช้ได้ แต่ละตัวอยู่ในรูป {"type": "function", "function": {"name", "description", "parameters"}} การเรียกของโมเดลจะส่งกลับมาใน tool_calls และโค้ดของคุณเป็นผู้รัน | ทุกโมเดล | |
tool_choice | string | object | auto | "auto" ให้โมเดลตัดสินใจเอง "required" บังคับให้เรียก tool {"type": "function", "function": {"name": "…"}} บังคับให้เรียก tool นั้น | โมเดล open-weight ที่โฮสต์ |
response_format | object | {"type": "json_object"} สำหรับคำตอบ JSON หรือ {"type": "json_schema", "json_schema": {…}} สำหรับคำตอบที่เป็นไปตาม schema ของคุณ | ทุกระดับของ Shannon; โมเดล open-weight ที่โฮสต์ตามที่ระบุไว้ในแต่ละ id | |
web_search | boolean | false | true ให้โมเดลค้นหาเว็บก่อนตอบ | shannon-1.6-*, shannon-2-*, ตระกูล Shannon 3 |
ฟิลด์อื่นของ OpenAI เช่น n, user, stream_options, parallel_tool_calls, presence_penalty, frequency_penalty, logit_bias, logprobs, metadata, store และ prompt_cache_key ระบบรับไว้เพื่อให้โค้ดไคลเอนต์เดิมทำงานได้โดยไม่ต้องแก้ ฟิลด์เหล่านี้ไม่เปลี่ยนคำตอบ: จะมี choice เดียวเสมอ และสตรีมจะจบด้วย usage เสมอ
ฟิลด์ที่มีชนิด JSON ผิด เช่น "max_tokens": "100" จะคืน 422 คำขอที่ไม่มี messages ก็เช่นกัน
Tools, เอาต์พุตแบบมีโครงสร้าง, การให้เหตุผล และการค้นหาเว็บ ต่างมีหน้าของตัวเอง: การเรียกฟังก์ชัน, เอาต์พุตแบบมีโครงสร้าง, ระดับความพยายามในการให้เหตุผล, ค้นหาเว็บในตัว.
คำขอที่มีตัวเลือก
คำขอนี้กำหนดข้อความ system, ฟิลด์ sampling และระดับความพยายามในการให้เหตุผล โดยใช้โมเดล open-weight ที่โฮสต์ ซึ่งรองรับครบทุกตัว
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.shannon-ai.com/v1",
)
response = client.chat.completions.create(
model="DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
messages=[
{"role": "system", "content": "You are a physics teacher. Answer in two sentences."},
{"role": "user", "content": "Why is the sky blue?"},
],
max_tokens=512,
temperature=0.3,
top_p=0.9,
seed=7,
stop=["\n\n"],
reasoning_effort="low",
)
message = response.choices[0].message
print(message.reasoning_content) # the reasoning
print(message.content) # the answer
print(response.usage) import OpenAI from "openai";
const client = new OpenAI({
apiKey: "YOUR_API_KEY",
baseURL: "https://api.shannon-ai.com/v1",
});
const response = await client.chat.completions.create({
model: "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
messages: [
{ role: "system", content: "You are a physics teacher. Answer in two sentences." },
{ role: "user", content: "Why is the sky blue?" },
],
max_tokens: 512,
temperature: 0.3,
top_p: 0.9,
seed: 7,
stop: ["\n\n"],
reasoning_effort: "low",
});
const message = response.choices[0].message;
console.log(message.reasoning_content); // the reasoning
console.log(message.content); // the answer
console.log(response.usage); curl https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP",
"messages": [
{"role": "system", "content": "You are a physics teacher. Answer in two sentences."},
{"role": "user", "content": "Why is the sky blue?"}
],
"max_tokens": 512,
"temperature": 0.3,
"top_p": 0.9,
"seed": 7,
"stop": ["\n\n"],
"reasoning_effort": "low"
}' คำตอบมีรูปแบบเดียวกับข้างต้น usage ของมันเพิ่มรายละเอียดสองอย่างบนโมเดล open-weight ที่โฮสต์: prompt tokens ที่อ่านจากแคช และ tokens ที่ใช้ไปกับการให้เหตุผล
{
"usage": {
"prompt_tokens": 31,
"completion_tokens": 62,
"total_tokens": 93,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 21
}
}
} ความยาวเอาต์พุต
max_tokens ทำสองอย่าง อย่างแรก คือจำนวน tokens ที่กันไว้จากยอดคงเหลือของคุณเมื่อคำขอเริ่ม เมื่อคำตอบสมบูรณ์ จำนวนนั้นจะถูกแทนที่ด้วย tokens ที่คำขอใช้จริง หาก max_tokens มากกว่ายอดที่เหลือในยอดคงเหลือของคุณ คำขอจะคืน 429 Quota exceeded แม้ว่าคำตอบจริงจะพอก็ตาม ส่ง max_tokens ที่ต่ำลงเพื่อกันไว้น้อยลง
shannon-coder-1 นับต่างออกไปบน endpoint นี้: แต่ละคำขอนับเป็นหนึ่งการเรียก Shannon Coder ของแผนของคุณ และไม่มีการกัน tokens ไว้ ขีดจำกัดและยอดคงเหลือ
อย่างที่สอง คือจำกัดความยาวของคำตอบบนโมเดลเหล่านี้:
| โมเดล | max_tokens ทำอะไร |
|---|---|
shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 | คำตอบหยุดเมื่อถึงขีดจำกัด จากนั้นสตรีมจะจบด้วย finish_reason length |
| โมเดล open-weight ที่โฮสต์ | ข้อความคำตอบจะหยุดที่ max_tokens การให้เหตุผลไม่นับรวม ค่าที่ต่ำกว่า 256 จะถูกใช้เป็น 256 |
หากไม่ระบุ max_tokens หรือ max_completion_tokens ค่าจะเป็น 4,096 และบน shannon-coder-1 เป็น 65,536
Messages
แต่ละข้อความเป็นออบเจกต์ที่มี role และ content โดย content เป็นสตริง หรืออาร์เรย์ของส่วนต่างๆ เมื่อข้อความมีมากกว่าข้อความธรรมดา
| Role | คำอธิบาย | ใช้โดย |
|---|---|---|
system | คำสั่งสำหรับโมเดล ใส่ไว้เป็นอันดับแรก บนระดับ Shannon จะใช้ข้อความ system ข้อความแรก | โมเดล open-weight ที่โฮสต์, shannon-1.6-*, shannon-2-*, shannon-coder-1 |
developer | อ่านเป็น system | โมเดล open-weight ที่โฮสต์ |
user | สิ่งที่คุณถาม บนระดับ Shannon ข้อความ user ล่าสุดคือ prompt และข้อความก่อนหน้าคือประวัติ | ทุกโมเดล |
assistant | คำตอบก่อนหน้าของโมเดล เก็บ tool_calls ของมันไว้เมื่อคุณส่งผลลัพธ์ของ tool ตามหลัง | ทุกโมเดล |
tool | ผลลัพธ์ของการเรียก tool: tool_call_id เก็บ id ของการเรียก และ content เก็บผลลัพธ์เป็นสตริง | ทุกโมเดล |
กับ id ในตระกูล Shannon 3 ให้ใส่คำสั่งที่ต้องคงไว้ลงในข้อความ user
บนระดับ Shannon คำขอที่ไม่มีข้อความของผู้ใช้และไม่มี tools จะคืน 400 No user message provided
ส่วนของเนื้อหา
| ส่วน | คำอธิบาย | ใช้ได้บน |
|---|---|---|
{"type": "text", "text": "…"} | ข้อความธรรมดา | ทุกโมเดล |
{"type": "image_url", "image_url": {"url": "…"}} | รูปภาพ ส่งเป็น URL แบบ data: ที่มีเนื้อหา base64 หรือเป็น URL แบบ http(s) | ตระกูล Shannon 3, shannon-1.6-lite, shannon-1.6-pro และโมเดล open-weight ที่โฮสต์ซึ่งระบุว่ารับอินพุตรูปภาพ |
{"type": "file", "source": {"type": "base64", "media_type": "application/pdf", "data": "…"}} | เอกสาร (PDF, Word, PowerPoint หรือ Excel) ส่งเป็น base64 หรือผ่าน URL | ตระกูล Shannon 3 |
ขนาด ข้อจำกัด และรายการรูปแบบทั้งหมดมีหน้าของตัวเอง รูปภาพและไฟล์
ออบเจกต์คำตอบ
| ฟิลด์ | ประเภท | คำอธิบาย |
|---|---|---|
id | string | chatcmpl- ตามด้วยอักขระเลขฐานสิบหก 32 ตัว |
object | string | เป็น chat.completion เสมอ |
created | integer | เวลาของคำตอบ เป็นวินาทีแบบ Unix |
model | string | id มาตรฐานของโมเดลที่ตอบ อาจสะกดต่างจาก id ที่คุณส่ง |
choices | array | มี choice เดียวเสมอ โดยมี index เป็น 0 |
choices[0].message.role | string | เป็น assistant เสมอ |
choices[0].message.content | string | null | ข้อความคำตอบ เมื่อมี tool_calls จะเป็น null บนระดับ Shannon ส่วนโมเดล open-weight ที่โฮสต์ส่งข้อความควบคู่ไปกับการเรียกได้ |
choices[0].message.reasoning_content | string | null | การให้เหตุผลที่โมเดลเขียนก่อนคำตอบ หรือ null เมื่อไม่มี |
choices[0].message.tool_calls | array | มีเฉพาะเมื่อโมเดลเรียก tools แต่ละรายการมี id, type เป็น function และ function ที่มี name กับ arguments เป็นสตริง JSON |
choices[0].message.annotations | array | มีเฉพาะในคำขอที่มี web_search: true และการค้นหาพบผลลัพธ์ หนึ่ง url_citation สำหรับแต่ละแหล่งที่มาที่เครื่องหมายใน content ระบุ ประกอบด้วย url, title, start_index และ end_index (ตำแหน่งของเครื่องหมาย นับเป็นตัวอักษร ไม่รวมตำแหน่งสิ้นสุด) |
choices[0].finish_reason | string | เหตุผลที่คำตอบจบ ดู เหตุผลที่จบ |
usage | object | tokens ของคำขอ ดู Usage |
sources | array | มีเฉพาะในคำขอที่มี web_search: true และการค้นหาพบผลลัพธ์: ผลลัพธ์ที่ส่งให้โมเดล แต่ละรายการมี index, title และ url [1] ในคำตอบคือรายการที่มี index เป็น 1 |
เหตุผลที่จบ
| finish_reason | คำอธิบาย |
|---|---|
stop | โมเดลตอบจบแล้ว หรือสตริง stop ปรากฏขึ้น |
tool_calls | โมเดลเรียก tool อย่างน้อยหนึ่งตัว ให้รัน tool เหล่านั้นแล้วส่งผลลัพธ์ในข้อความ tool |
length | คำตอบถูกตัดที่ขีดจำกัดของเอาต์พุต รายงานในสตรีมของ shannon-1.6-lite, shannon-1.6-pro, shannon-coder-1 และตระกูล Shannon 3 |
คำตอบที่ไม่ได้สตรีมจะรายงาน stop หรือ tool_calls
Usage
| ฟิลด์ | ประเภท | คำอธิบาย | ใช้ได้บน |
|---|---|---|---|
usage.prompt_tokens | integer | Input tokens | ทุกโมเดล |
usage.completion_tokens | integer | Output tokens: การให้เหตุผล คำตอบ และการเรียก tool รวมกัน | ทุกโมเดล |
usage.total_tokens | integer | prompt_tokens บวก completion_tokens | ทุกโมเดล |
usage.prompt_tokens_details.cached_tokens | integer | ส่วนของ prompt_tokens ที่อ่านจาก prompt cache | โมเดล open-weight ที่โฮสต์ |
usage.completion_tokens_details.reasoning_tokens | integer | ส่วนของ completion_tokens ที่ใช้ไปกับการให้เหตุผล | โมเดล open-weight ที่โฮสต์ |
บนโมเดล open-weight ที่โฮสต์ prompt_tokens คือข้อความและนิยาม tool ของคุณที่นับด้วย tokenizer ของโมเดลเอง บวก tokens ของรูปภาพ (ถ้ามี) endpoint นับ tokens คืนตัวเลขเดียวกันก่อนที่คุณจะส่ง การนับ token
บนระดับ Shannon prompt_tokens นับทุกอย่างที่โมเดลอ่านเพื่อเขียนคำตอบ จึงมากกว่าข้อความของคุณเพียงอย่างเดียว
การสตรีม
เมื่อตั้งค่า stream เป็น true คำตอบจะมาถึงเป็นอีเวนต์ chat.completion.chunk และจบด้วย data: [DONE] chunk สุดท้ายก่อนหน้านั้นมี finish_reason และ usage ไม่ต้องใช้ stream_options รูปแบบของ chunk บรรทัด keep-alive และข้อผิดพลาดภายในสตรีมมีหน้าของตัวเอง สตรีมมิง
ข้อผิดพลาด
ข้อผิดพลาดคือออบเจกต์ JSON ที่มีสมาชิก error การตรวจสอบทำตามลำดับนี้: API key, body ของคำขอ, model id แล้วจึงยอดคงเหลือ ตารางแสดงสิ่งที่ endpoint นี้คืนบ่อยที่สุด รายการทั้งหมดพร้อมวิธีตัดสินใจว่าควรลองใหม่หรือไม่มีหน้าของตัวเอง การจัดการข้อผิดพลาด
{
"error": {
"type": "invalid_request_error",
"message": "unknown model: no-such-model"
}
} | สถานะ | ประเภท | ข้อความ | เมื่อไร |
|---|---|---|---|
401 | authentication_error | Missing authenticationInvalid API key | ไม่ได้ส่ง API key มา หรือคีย์ไม่รู้จักหรือถูกเพิกถอนแล้ว |
400 | invalid_request_error | unknown model: <id> | model ไม่ใช่ id ที่เผยแพร่ |
400 | invalid_request_error | No user message provided | ระดับ Shannon: คำขอไม่มีข้อความของผู้ใช้และไม่มี tools |
400 | invalid_request_error | <id> does not accept image input | ส่งส่วนรูปภาพไปยังโมเดล open-weight ที่โฮสต์ซึ่งไม่รองรับอินพุตรูปภาพ |
400 | invalid_request_error | <id> does not accept response_format | ส่ง response_format ไปยังโมเดล open-weight ที่โฮสต์ซึ่งไม่รองรับเอาต์พุตแบบมีโครงสร้าง |
400 | invalid_request_error | unknown reasoning effort '<value>'; expected off, low, medium or high | reasoning_effort มีค่าที่อยู่นอกรายการ |
422 | invalid_request_error | Failed to deserialize the JSON body into the target type: … | ไม่มี messages หรือฟิลด์ใดฟิลด์หนึ่งมีชนิด JSON ผิด |
429 | rate_limit_error | Quota exceeded. Upgrade your plan at shannon-ai.com/plan | max_tokens มากกว่ายอดที่เหลือในยอดคงเหลือของคุณ |
429 | rate_limit_error | Too many requests. Retry in <n>s. | Flood protection: มีคำขอมากกว่า 120 รายการในหนึ่งนาทีบนบัญชีของคุณ |
500 | server_error | The model backend failed to answer. Please retry. | โมเดลไม่ได้สร้างคำตอบ ส่งคำขอใหม่อีกครั้ง |
502 | api_error | The model backend failed to answer. Please retry. | เช่นเดียวกัน บนตระกูล Shannon 3 และโมเดล open-weight ที่โฮสต์ |