Images and files
Send images with a message: which models take them, the accepted forms and their limits.
POST https://api.shannon-ai.com/v1/chat/completions
An image travels inside a user message, as one part of its content. The sample reads a local file and sends it as a data URL.
import base64
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")
with open("photo.jpg", "rb") as f:
image = base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
model="shannon-3",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64," + image}},
],
}],
)
print(response.choices[0].message.content) import fs from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({ apiKey: "YOUR_API_KEY", baseURL: "https://api.shannon-ai.com/v1" });
const image = fs.readFileSync("photo.jpg").toString("base64");
const response = await client.chat.completions.create({
model: "shannon-3",
messages: [{
role: "user",
content: [
{ type: "text", text: "What is in this picture?" },
{ type: "image_url", image_url: { url: "data:image/jpeg;base64," + image } },
],
}],
});
console.log(response.choices[0].message.content); # Linux: base64 -w0 photo.jpg macOS: base64 -i photo.jpg
IMAGE=$(base64 -w0 photo.jpg)
curl https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<EOF
{
"model": "shannon-3",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,${IMAGE}"}}
]
}]
}
EOF {
"id": "chatcmpl-8d1f3a5c7e9b4d2f6a8c0e2b4d6f8a1c",
"object": "chat.completion",
"created": 1791590400,
"model": "shannon-3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "A grey cat asleep on a windowsill, with a potted fern beside it.",
"reasoning_content": null
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 286,
"completion_tokens": 19,
"total_tokens": 305
}
} Which models take images
| Models | Reads |
|---|---|
shannon-3, shannon-3-pro, shannon-3.1, shannon-3.1-pro | Images, and documents: PDF, Word, PowerPoint, Excel and plain-text files. |
shannon-1.6-lite, shannon-1.6-pro | Images. |
shannon-2-lite, shannon-2-pro | Images, in requests that also send tools or response_format. |
Kimi-K3-3BIT-REAP, MiniMax-M3-3BIT-REAP, Kimi-K2.6-W4A16-AUTOROUND-REAP, inkling-W4A16-AUTOROUND-REAP, MiMo-V2.5-W8A16 | Images. |
DeepSeek-V4-Pro-0813-3BIT-REAP, GLM-5.2-3BIT-REAP, Nemotron3Ultra-3BIT-REAP, DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP, Laguna-S-2.1-W4A16-AUTOROUND-REAP, MiMo-V2.5-Pro-W8A16, Hy3-W8A16 | Text. A request with an image is answered with status 400. |
shannon-coder-1 | Text. |
GET /v1/models reports image input per id as capabilities.vision. Models & pricing
- The Shannon models read the files of the last user message. Put the image in the message that asks about it.
- The Shannon 3 family also keeps a few images and documents of earlier user messages in view when they were sent inline, and reads a set maximum of images per request, counted from the first.
- The hosted open-weight models read the images of every message of the conversation.
Accepted forms
A file is sent in one of two ways: inside the request as base64, or as an address the API fetches. Every file travels with the request; the API has no upload endpoint.
| Endpoint | In the request (base64) | By address |
|---|---|---|
/v1/chat/completions | {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,…"}} | {"type": "image_url", "image_url": {"url": "https://…"}} |
/v1/responses | {"type": "input_image", "image_url": "data:image/jpeg;base64,…"} | {"type": "input_image", "image_url": "https://…"} |
/v1/messages | {"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": "…"}} | {"type": "image", "source": {"type": "url", "url": "https://…"}} |
The same user message in each format:
{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{
"type": "image_url",
"image_url": {
"url": "data:image/jpeg;base64,/9j/4AAQSkZJRgABAQAAAQABAAD…"
}
}
]
} {
"role": "user",
"content": [
{"type": "input_text", "text": "What is in this picture?"},
{
"type": "input_image",
"image_url": "data:image/jpeg;base64,/9j/4AAQSkZJRgABAQAAAQABAAD…"
}
]
} {
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/jpeg",
"data": "/9j/4AAQSkZJRgABAQAAAQABAAD…"
}
}
]
} - A data URL has the form
data:<media type>;base64,<data>. The;base64part is required. - On
/v1/chat/completionsand/v1/responses,image_urlcan be the string itself or an object{"url": "…"}. - The file type comes from the data URL or from
media_type. Without one, an image is read asimage/png. - Inline files count toward the size of the request body, which can be up to 32 MiB.
- Accepted for compatibility:
image_url.detail.
Limits for remote files
When you send an address, the API fetches the file before the model runs, within these limits:
| Limit | Rule |
|---|---|
| Size | Up to 8 MiB per file. |
| Time | 20 seconds to answer. |
| Redirects | At most 5. |
| Address | http:// or https://, without a user name or password in the address, on a host with a public address. |
| Reply of the host | A status in the 200 range and a body that is not empty. The file type is read from Content-Type. |
| Requested as | User-Agent: ShannonBot/1.0 (+https://shannon-ai.com) |
How images count in tokens
On the hosted open-weight models an image counts one token for every 28 × 28 pixel patch: the width divided by 28, rounded up, times the height divided by 28, rounded up. The result is part of usage.prompt_tokens and is billed as input.
- 1024 × 1024
- 1,369 tokens
- 1920 × 1080
- 2,691 tokens
- 512 × 512
- 361 tokens
- A remote image is fetched first, so its real size is counted.
- An image whose size cannot be read counts 1,024 tokens.
- The counting endpoints use the same rule for data URLs. They do not fetch a remote address and count it as 1,024 tokens. Token counting
- For a Shannon model, read what a request with images cost from
usagein the reply.
Documents and other files
The Shannon 3 family (shannon-3, shannon-3-pro, shannon-3.1, shannon-3.1-pro) reads documents as well. The text of the document is taken out and given to the model with your message.
| Document | Media type |
|---|---|
application/pdf | |
| Word | application/vnd.openxmlformats-officedocument.wordprocessingml.document |
| PowerPoint | application/vnd.openxmlformats-officedocument.presentationml.presentation |
| Excel | application/vnd.openxmlformats-officedocument.spreadsheetml.sheet |
A document is a part of the user message, like an image:
| Endpoint | In the request (base64) | By address |
|---|---|---|
/v1/chat/completions, /v1/messages | {"type": "document", "source": {"type": "base64", "media_type": "application/pdf", "data": "…"}} | {"type": "document", "source": {"type": "url", "url": "https://…"}} |
/v1/responses | {"type": "input_file", "file_data": "data:application/pdf;base64,…"} | {"type": "input_file", "file_url": "https://…"} |
{
"model": "shannon-3",
"messages": [
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": "JVBERi0xLjcKJeLjz9MK…"
}
},
{"type": "text", "text": "Summarise this report in five points."}
]
}
]
} {
"model": "shannon-3",
"input": [
{
"role": "user",
"content": [
{
"type": "input_file",
"file_data": "data:application/pdf;base64,JVBERi0xLjcKJeLjz9MK…"
},
{
"type": "input_text",
"text": "Summarise this report in five points."
}
]
}
]
} {
"model": "shannon-3",
"messages": [
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": "JVBERi0xLjcKJeLjz9MK…"
}
},
{"type": "text", "text": "Summarise this report in five points."}
]
}
]
} - Give the real media type of the file. Without one, a document is read as
application/pdf. - A plain-text file, such as CSV, is read as text and attached in the same way.
- When a document cannot be read, or the file is of another type, the model is told so and can say it in its answer. The status is still
200. - On every other model a document part is left out of what the model reads, and the request is answered from the rest.
- A document sent by address is fetched under the limits for remote files above.
Errors
| Status | Type | Message | When |
|---|---|---|---|
400 | invalid_request_error | <id> does not accept image input | An image was sent to a hosted open-weight model that takes text only. Nothing is drawn from your balance. |
400 | invalid_request_error | audio and video input are not supported by any hosted model | An audio or video part was sent to a hosted open-weight model. |
413 | invalid_request_error | Failed to buffer the request body: … | The request body, with its inline files, is larger than 32 MiB. The error carries code: "request_too_large". |
A Shannon model answers with status 200 whatever files the request carries: a part it does not read is left out. Error Handling