图片与文件
随消息发送图片:哪些模型支持、接受的形式及其限制。
POST https://api.shannon-ai.com/v1/chat/completions
图片放在用户消息中传送,作为其 content 的一个部分。示例读取本地文件,并以 data URL 形式发送。
import base64
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")
with open("photo.jpg", "rb") as f:
image = base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
model="shannon-3",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64," + image}},
],
}],
)
print(response.choices[0].message.content) import fs from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({ apiKey: "YOUR_API_KEY", baseURL: "https://api.shannon-ai.com/v1" });
const image = fs.readFileSync("photo.jpg").toString("base64");
const response = await client.chat.completions.create({
model: "shannon-3",
messages: [{
role: "user",
content: [
{ type: "text", text: "What is in this picture?" },
{ type: "image_url", image_url: { url: "data:image/jpeg;base64," + image } },
],
}],
});
console.log(response.choices[0].message.content); # Linux: base64 -w0 photo.jpg macOS: base64 -i photo.jpg
IMAGE=$(base64 -w0 photo.jpg)
curl https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<EOF
{
"model": "shannon-3",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,${IMAGE}"}}
]
}]
}
EOF 200 JSON
{
"id": "chatcmpl-8d1f3a5c7e9b4d2f6a8c0e2b4d6f8a1c",
"object": "chat.completion",
"created": 1791590400,
"model": "shannon-3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "A grey cat asleep on a windowsill, with a potted fern beside it.",
"reasoning_content": null
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 286,
"completion_tokens": 19,
"total_tokens": 305
}
} 哪些模型支持图片
| 模型 | 读取 |
|---|---|
shannon-3, shannon-3-pro, shannon-3.1, shannon-3.1-pro | 图片和文档:PDF、Word、PowerPoint、Excel 和纯文本文件。 |
shannon-1.6-lite, shannon-1.6-pro | 图片。 |
shannon-2-lite, shannon-2-pro | 图片,限于同时发送 tools 或 response_format 的请求。 |
Kimi-K3-3BIT-REAP, MiniMax-M3-3BIT-REAP, Kimi-K2.6-W4A16-AUTOROUND-REAP, inkling-W4A16-AUTOROUND-REAP, MiMo-V2.5-W8A16 | 图片。 |
DeepSeek-V4-Pro-0813-3BIT-REAP, GLM-5.2-3BIT-REAP, Nemotron3Ultra-3BIT-REAP, DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP, Laguna-S-2.1-W4A16-AUTOROUND-REAP, MiMo-V2.5-Pro-W8A16, Hy3-W8A16 | 文本。带图片的请求会得到状态码 400。 |
shannon-coder-1 | 文本。 |
GET /v1/models 以 capabilities.vision 报告每个 id 的图片输入。 模型与定价
- Shannon 模型读取最后一条用户消息中的文件。请把图片放在提问的那条消息里。
- 对于以内联方式发送的内容,Shannon 3 系列还会保留较早用户消息中的少量图片和文档,并且每个请求读取的图片数量有固定上限,从第一张开始计数。
- 托管开源权重模型会读取对话中每条消息的图片。
接受的形式
文件有两种发送方式:以 base64 放在请求内,或提供一个由 API 去获取的地址。每个文件都随请求一起传送;API 没有上传端点。
| 端点 | 在请求中(base64) | 通过地址 |
|---|---|---|
/v1/chat/completions | {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,…"}} | {"type": "image_url", "image_url": {"url": "https://…"}} |
/v1/responses | {"type": "input_image", "image_url": "data:image/jpeg;base64,…"} | {"type": "input_image", "image_url": "https://…"} |
/v1/messages | {"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": "…"}} | {"type": "image", "source": {"type": "url", "url": "https://…"}} |
各格式中的同一条用户消息:
{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{
"type": "image_url",
"image_url": {
"url": "data:image/jpeg;base64,/9j/4AAQSkZJRgABAQAAAQABAAD…"
}
}
]
} {
"role": "user",
"content": [
{"type": "input_text", "text": "What is in this picture?"},
{
"type": "input_image",
"image_url": "data:image/jpeg;base64,/9j/4AAQSkZJRgABAQAAAQABAAD…"
}
]
} {
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/jpeg",
"data": "/9j/4AAQSkZJRgABAQAAAQABAAD…"
}
}
]
} - data URL 的形式为
data:<media type>;base64,<data>。;base64部分是必需的。 - 在
/v1/chat/completions和/v1/responses上,image_url可以是字符串本身,也可以是对象{"url": "…"}。 - 文件类型取自 data URL 或
media_type。如果没有,图片会按image/png读取。 - 内联文件计入请求体的大小,请求体最大可达 32 MiB。
- 为兼容而接受:
image_url.detail。
远程文件的限制
当你发送地址时,API 会在模型运行之前获取该文件,并遵循以下限制:
| 限制 | 规则 |
|---|---|
| 大小 | 每个文件最大 8 MiB。 |
| 时间 | 20 秒内须作出答复。 |
| 重定向 | 最多 5 次。 |
| 地址 | http:// 或 https://,地址中不含用户名或密码,主机须有公网地址。 |
| 主机的回复 | 200 范围内的状态码,且响应体不为空。文件类型取自 Content-Type。 |
| 请求标识为 | User-Agent: ShannonBot/1.0 (+https://shannon-ai.com) |
图片如何折算为 tokens
在托管开源权重模型上,每个 28 × 28 像素的图块计为一个 token:宽度除以 28 向上取整,乘以高度除以 28 向上取整。结果是 usage.prompt_tokens 的一部分,并按输入计费。
- 1024 × 1024
- 1,369 tokens
- 1920 × 1080
- 2,691 tokens
- 512 × 512
- 361 tokens
- 远程图片会先被获取,因此按其真实大小计算。
- 无法读取大小的图片按 1,024 tokens 计算。
- 计数端点对 data URL 使用相同的规则。它们不会获取远程地址,并将其按 1,024 tokens 计算。 Token 计数
- 对于 Shannon 模型,请从回复的
usage中查看带图片的请求花费了多少。
文档及其他文件
Shannon 3 系列(shannon-3、shannon-3-pro、shannon-3.1、shannon-3.1-pro)也能读取文档。文档中的文本会被提取出来,与你的消息一起交给模型。
| 文档 | 媒体类型 |
|---|---|
application/pdf | |
| Word | application/vnd.openxmlformats-officedocument.wordprocessingml.document |
| PowerPoint | application/vnd.openxmlformats-officedocument.presentationml.presentation |
| Excel | application/vnd.openxmlformats-officedocument.spreadsheetml.sheet |
文档与图片一样,是用户消息的一部分:
| 端点 | 在请求中(base64) | 通过地址 |
|---|---|---|
/v1/chat/completions, /v1/messages | {"type": "document", "source": {"type": "base64", "media_type": "application/pdf", "data": "…"}} | {"type": "document", "source": {"type": "url", "url": "https://…"}} |
/v1/responses | {"type": "input_file", "file_data": "data:application/pdf;base64,…"} | {"type": "input_file", "file_url": "https://…"} |
{
"model": "shannon-3",
"messages": [
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": "JVBERi0xLjcKJeLjz9MK…"
}
},
{"type": "text", "text": "Summarise this report in five points."}
]
}
]
} {
"model": "shannon-3",
"input": [
{
"role": "user",
"content": [
{
"type": "input_file",
"file_data": "data:application/pdf;base64,JVBERi0xLjcKJeLjz9MK…"
},
{
"type": "input_text",
"text": "Summarise this report in five points."
}
]
}
]
} {
"model": "shannon-3",
"messages": [
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": "JVBERi0xLjcKJeLjz9MK…"
}
},
{"type": "text", "text": "Summarise this report in five points."}
]
}
]
} - 请给出文件真实的媒体类型。如果未给出,文档会按
application/pdf读取。 - 纯文本文件(如 CSV)会按文本读取,并以同样方式附加。
- 当文档无法读取,或文件属于其他类型时,模型会被告知这一点,并可在答案中说明。状态码仍为
200。 - 在其他所有模型上,文档部分不会出现在模型所读取的内容中,请求将根据其余部分作答。
- 通过地址发送的文档,在获取时适用上文针对远程文件的限制。
错误
| 状态码 | 类型 | 消息 | 何时 |
|---|---|---|---|
400 | invalid_request_error | <id> does not accept image input | 向只接受文本的托管开源权重模型发送了图片。不会从你的余额中扣除任何内容。 |
400 | invalid_request_error | audio and video input are not supported by any hosted model | 向托管开源权重模型发送了音频或视频部分。 |
413 | invalid_request_error | Failed to buffer the request body: … | 请求体连同其内联文件大于 32 MiB。该错误带有 code: "request_too_large"。 |
无论请求携带什么文件,Shannon 模型都返回状态码 200:它不读取的部分会被忽略。 错误处理