이미지와 파일
메시지와 함께 이미지 보내기: 어떤 모델이 받는지, 허용되는 형식, 제한.
POST https://api.shannon-ai.com/v1/chat/completions
이미지는 사용자 메시지 안에서 content의 한 파트로 전달됩니다. 샘플은 로컬 파일을 읽어 data URL로 보냅니다.
import base64
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")
with open("photo.jpg", "rb") as f:
image = base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
model="shannon-3",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64," + image}},
],
}],
)
print(response.choices[0].message.content) import fs from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({ apiKey: "YOUR_API_KEY", baseURL: "https://api.shannon-ai.com/v1" });
const image = fs.readFileSync("photo.jpg").toString("base64");
const response = await client.chat.completions.create({
model: "shannon-3",
messages: [{
role: "user",
content: [
{ type: "text", text: "What is in this picture?" },
{ type: "image_url", image_url: { url: "data:image/jpeg;base64," + image } },
],
}],
});
console.log(response.choices[0].message.content); # Linux: base64 -w0 photo.jpg macOS: base64 -i photo.jpg
IMAGE=$(base64 -w0 photo.jpg)
curl https://api.shannon-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<EOF
{
"model": "shannon-3",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,${IMAGE}"}}
]
}]
}
EOF {
"id": "chatcmpl-8d1f3a5c7e9b4d2f6a8c0e2b4d6f8a1c",
"object": "chat.completion",
"created": 1791590400,
"model": "shannon-3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "A grey cat asleep on a windowsill, with a potted fern beside it.",
"reasoning_content": null
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 286,
"completion_tokens": 19,
"total_tokens": 305
}
} 이미지를 받는 모델
| 모델 | 읽는 것 |
|---|---|
shannon-3, shannon-3-pro, shannon-3.1, shannon-3.1-pro | 이미지와 문서: PDF, Word, PowerPoint, Excel, 일반 텍스트 파일. |
shannon-1.6-lite, shannon-1.6-pro | 이미지. |
shannon-2-lite, shannon-2-pro | tools 또는 response_format도 함께 보내는 요청에서 이미지. |
Kimi-K3-3BIT-REAP, MiniMax-M3-3BIT-REAP, Kimi-K2.6-W4A16-AUTOROUND-REAP, inkling-W4A16-AUTOROUND-REAP, MiMo-V2.5-W8A16 | 이미지. |
DeepSeek-V4-Pro-0813-3BIT-REAP, GLM-5.2-3BIT-REAP, Nemotron3Ultra-3BIT-REAP, DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP, Laguna-S-2.1-W4A16-AUTOROUND-REAP, MiMo-V2.5-Pro-W8A16, Hy3-W8A16 | 텍스트. 이미지가 있는 요청에는 상태 400으로 응답합니다. |
shannon-coder-1 | 텍스트. |
GET /v1/models는 id별 이미지 입력을 capabilities.vision으로 보고합니다. 모델 및 가격
- Shannon 모델은 마지막 사용자 메시지의 파일을 읽습니다. 이미지는 그것에 대해 묻는 메시지에 넣으세요.
- Shannon 3 제품군은 이전 사용자 메시지의 이미지와 문서도 인라인으로 보냈다면 일부 유지하며, 요청당 읽는 이미지 수는 첫 번째부터 세어 정해진 최대치까지입니다.
- 호스팅 오픈 웨이트 모델은 대화의 모든 메시지에 있는 이미지를 읽습니다.
허용되는 형식
파일은 두 가지 방식 중 하나로 보냅니다: base64로 요청 안에 넣거나, API가 가져가는 주소로 보냅니다. 모든 파일은 요청과 함께 전달되며, API에는 업로드 엔드포인트가 없습니다.
| 엔드포인트 | 요청 내(base64) | 주소로 |
|---|---|---|
/v1/chat/completions | {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,…"}} | {"type": "image_url", "image_url": {"url": "https://…"}} |
/v1/responses | {"type": "input_image", "image_url": "data:image/jpeg;base64,…"} | {"type": "input_image", "image_url": "https://…"} |
/v1/messages | {"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": "…"}} | {"type": "image", "source": {"type": "url", "url": "https://…"}} |
각 형식의 동일한 사용자 메시지:
{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{
"type": "image_url",
"image_url": {
"url": "data:image/jpeg;base64,/9j/4AAQSkZJRgABAQAAAQABAAD…"
}
}
]
} {
"role": "user",
"content": [
{"type": "input_text", "text": "What is in this picture?"},
{
"type": "input_image",
"image_url": "data:image/jpeg;base64,/9j/4AAQSkZJRgABAQAAAQABAAD…"
}
]
} {
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/jpeg",
"data": "/9j/4AAQSkZJRgABAQAAAQABAAD…"
}
}
]
} - data URL의 형식은
data:<media type>;base64,<data>입니다.;base64부분은 필수입니다. /v1/chat/completions와/v1/responses에서image_url은 문자열 자체이거나{"url": "…"}객체일 수 있습니다.- 파일 유형은 data URL 또는
media_type에서 가져옵니다. 지정하지 않으면 이미지는image/png로 읽습니다. - 인라인 파일은 요청 본문 크기에 포함되며, 요청 본문은 최대 32 MiB까지 가능합니다.
- 호환성을 위해 허용:
image_url.detail.
원격 파일 제한
주소를 보내면 API가 모델을 실행하기 전에 다음 제한 안에서 파일을 가져옵니다:
| 제한 | 규칙 |
|---|---|
| 크기 | 파일당 최대 8 MiB. |
| 시간 | 응답까지 20초. |
| 리디렉션 | 최대 5회. |
| 주소 | http:// 또는 https://이며, 주소에 사용자 이름이나 비밀번호가 없어야 하고, 공개 주소를 가진 호스트여야 합니다. |
| 호스트의 응답 | 200번대 상태와 비어 있지 않은 본문. 파일 유형은 Content-Type에서 읽습니다. |
| 요청 방식 | User-Agent: ShannonBot/1.0 (+https://shannon-ai.com) |
이미지가 토큰으로 계산되는 방식
호스팅 오픈 웨이트 모델에서 이미지는 28 × 28 픽셀 패치마다 토큰 하나로 계산합니다: 너비를 28로 나눠 올림한 값에 높이를 28로 나눠 올림한 값을 곱합니다. 결과는 usage.prompt_tokens에 포함되며 입력으로 청구됩니다.
- 1024 × 1024
- 1,369 토큰
- 1920 × 1080
- 2,691 토큰
- 512 × 512
- 361 토큰
- 원격 이미지는 먼저 가져오므로 실제 크기로 계산됩니다.
- 크기를 읽을 수 없는 이미지는 1,024 토큰으로 계산합니다.
- 토큰 계산 엔드포인트는 data URL에 같은 규칙을 사용합니다. 원격 주소는 가져오지 않고 1,024 토큰으로 계산합니다. 토큰 수 계산
- Shannon 모델에서 이미지가 있는 요청의 비용은 응답의
usage에서 확인하세요.
문서 및 기타 파일
Shannon 3 제품군(shannon-3, shannon-3-pro, shannon-3.1, shannon-3.1-pro)은 문서도 읽습니다. 문서의 텍스트를 추출하여 메시지와 함께 모델에 전달합니다.
| 문서 | 미디어 유형 |
|---|---|
application/pdf | |
| Word | application/vnd.openxmlformats-officedocument.wordprocessingml.document |
| PowerPoint | application/vnd.openxmlformats-officedocument.presentationml.presentation |
| Excel | application/vnd.openxmlformats-officedocument.spreadsheetml.sheet |
문서는 이미지와 마찬가지로 사용자 메시지의 한 파트입니다:
| 엔드포인트 | 요청 내(base64) | 주소로 |
|---|---|---|
/v1/chat/completions, /v1/messages | {"type": "document", "source": {"type": "base64", "media_type": "application/pdf", "data": "…"}} | {"type": "document", "source": {"type": "url", "url": "https://…"}} |
/v1/responses | {"type": "input_file", "file_data": "data:application/pdf;base64,…"} | {"type": "input_file", "file_url": "https://…"} |
{
"model": "shannon-3",
"messages": [
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": "JVBERi0xLjcKJeLjz9MK…"
}
},
{"type": "text", "text": "Summarise this report in five points."}
]
}
]
} {
"model": "shannon-3",
"input": [
{
"role": "user",
"content": [
{
"type": "input_file",
"file_data": "data:application/pdf;base64,JVBERi0xLjcKJeLjz9MK…"
},
{
"type": "input_text",
"text": "Summarise this report in five points."
}
]
}
]
} {
"model": "shannon-3",
"messages": [
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": "JVBERi0xLjcKJeLjz9MK…"
}
},
{"type": "text", "text": "Summarise this report in five points."}
]
}
]
} - 파일의 실제 미디어 유형을 지정하세요. 지정하지 않으면 문서는
application/pdf로 읽습니다. - CSV 같은 일반 텍스트 파일은 텍스트로 읽어 같은 방식으로 첨부합니다.
- 문서를 읽을 수 없거나 파일이 다른 유형이면 모델에 그 사실이 전달되며, 모델이 답변에서 이를 말할 수 있습니다. 상태는 여전히
200입니다. - 다른 모든 모델에서는 문서 파트가 모델이 읽는 내용에서 제외되고, 나머지 내용으로 요청에 응답합니다.
- 주소로 보낸 문서는 위의 원격 파일 제한에 따라 가져옵니다.
오류
| 상태 | 유형 | 메시지 | 발생 시점 |
|---|---|---|---|
400 | invalid_request_error | <id> does not accept image input | 텍스트만 받는 호스팅 오픈 웨이트 모델에 이미지를 보냈습니다. 잔액에서는 아무것도 차감되지 않습니다. |
400 | invalid_request_error | audio and video input are not supported by any hosted model | 호스팅 오픈 웨이트 모델에 오디오 또는 비디오 파트를 보냈습니다. |
413 | invalid_request_error | Failed to buffer the request body: … | 인라인 파일을 포함한 요청 본문이 32 MiB보다 큽니다. 오류에 code: "request_too_large"가 담겨 있습니다. |
Shannon 모델은 요청에 어떤 파일이 있든 상태 200으로 응답합니다. 읽지 않는 파트는 제외됩니다. 오류 처리