Limits and balance
Every request is served equally. No rate tiers. No separate API quota. You already paid for your tokens — use them as fast as you like.
This page explains what your balance is made of, what one request reserves and costs, how many requests you may send, and the few limits a single request can meet.
- value of 1M tokens of balance
- $5.00
- daily allowance renews
- 00:00 UTC
- flood protection, per account
- 120 requests / min
How requests are served
- No rate tiers — One rule limits how fast requests may arrive, and it is the same for every account and every plan: 120 requests per minute. There is no limit on tokens per minute.
- No separate API quota — The API spends the same balance as chat. A plan sets the size of today's allowance. It does not set a request rate.
- As fast as you like — Requests sent in parallel are accepted and wait in line. They are not refused for being parallel.
Your balance
Your balance is counted in tokens. 1,000,000 tokens of balance are worth $5.00, and every price on the Models & pricing page is a rate against that value.
At any moment the balance is the sum of two parts.
- Today's plan allowance — A number of tokens set by your plan. It is new every day at 00:00 UTC. What is left at the end of a day does not carry over.
- Purchased credit — Tokens you bought as a pack. Credit does not expire and works on every plan, including Free.
| Plan | Tokens per day | Worth |
|---|---|---|
| Free | 30,000 | $0.15 |
| Plus | 80,000 | $0.40 |
| Standard | 265,000 | $1.325 |
| Pro | 665,000 | $3.325 |
- Order of spending — Every request spends today's plan allowance first. Purchased credit is used only for what goes beyond the allowance on that day.
- Chat and API share it — There is one balance per account. An API key spends from the balance of the account that owns it, at the same prices as chat.
- Packs — Credit is sold in packs of 1,000,000 ($5.00), 2,000,000 ($10.00) and 5,000,000 ($25.00) tokens, or as an amount of your choice from 1,000,000 to 100,000,000 tokens at $5.00 per 1,000,000.
What a request reserves and what it costs
- Reserve — When a request arrives it reserves its output budget from your balance:
max_tokenson/v1/chat/completionsand/v1/messages,max_output_tokenson/v1/responses./v1/chat/completionsalso readsmax_completion_tokens. The default is 4,096 and the range is 1 to 65,536. - Admit — The request is accepted only if the reservation fits into what is left of your balance. A balance that is above zero but smaller than the output budget gets the
Quota exceededreply. Send a smallermax_tokensto use the rest. - Settle — When the answer is complete, the reservation is replaced by the real charge. The charge can be lower or higher than the reservation.
- Return — A request that ends with an error status gives its reservation back in full.
The real charge depends on the model family.
| Models | What is charged |
|---|---|
| Shannon models | usage.total_tokens at the model's price per 1M. Input and output have one rate. |
| Hosted open-weight models | Uncached input at the input rate, cached input at the cached rate, output at the output rate. |
The amount in USD is taken from your balance in tokens at $5.00 per 1,000,000, rounded to a whole token.
Counting tokens with POST /v1/tokenize or POST /v1/messages/count_tokens is free and reserves nothing. Token counting
Where to see balance and usage
The Keys & usage page shows what you can spend right now, today's plan allowance, your purchased credit and the API spend of the last 30 days. Below that it lists every request your key made: time, endpoint, model, cached input, billed tokens and cost. Keys & usage
Every reply also carries a usage object with the token counts of that call.
| Endpoint | Fields of usage | Added by hosted open-weight models |
|---|---|---|
/v1/chat/completions | prompt_tokens, completion_tokens, total_tokens | prompt_tokens_details.cached_tokens, completion_tokens_details.reasoning_tokens |
/v1/messages | input_tokens, output_tokens | cache_read_input_tokens, cache_creation_input_tokens |
/v1/responses | input_tokens, output_tokens, total_tokens | input_tokens_details.cached_tokens, output_tokens_details.reasoning_tokens |
usageholds the token counts of the model. The amount taken from your balance is not in the reply: it is the Billed tokens column of the request list on Keys & usage.- On
/v1/messageswith a hosted open-weight model,input_tokensis the uncached part of the input,cache_read_input_tokensis the cached part andcache_creation_input_tokensis always0. - A stream on
/v1/chat/completionscarriesusagein its last chunk before[DONE]. Streaming
When the balance runs out
A request whose reservation does not fit into your balance is answered with status 429, type rate_limit_error and the message below. Nothing is charged. The same reply is sent when the balance is above zero but smaller than the output budget of the request.
{
"error": {
"type": "rate_limit_error",
"message": "Quota exceeded. Upgrade your plan at shannon-ai.com/plan"
}
} {
"type": "error",
"error": {
"type": "rate_limit_error",
"message": "Quota exceeded. Upgrade your plan at shannon-ai.com/plan"
}
} On /v1/responses the error object can also hold code and param, both null.
What you can do:
- Wait for the next plan allowance at 00:00 UTC.
- Top up credit. Credit is spent after the plan allowance and does not expire. Top up credit
- Change to a plan with a larger daily allowance. Change plan
- Send a smaller
max_tokens, if some balance is left: the reservation is then smaller.
Shannon Coder call allowance
shannon-coder-1 on /v1/chat/completions and /v1/messages is counted in calls, not in tokens. Each plan includes a number of calls per 4-hour window. One request is one call.
| Plan | Calls per 4-hour window |
|---|---|
| Free | 3 |
| Plus | 20 |
| Standard | 40 |
| Pro | 60 |
- Windows start at 00:00, 04:00, 08:00, 12:00, 16:00, 20:00 UTC. Calls that are left at the end of a window do not carry over.
- A call is counted when the request is accepted, before the model answers. A request that fails afterwards still counts as a call.
- These calls reserve no tokens and take nothing from your balance. The request list on Keys & usage shows their token count and its value at the listed price.
- The default
max_tokensofshannon-coder-1on these two endpoints is 65,536. - With no calls left the reply is status
429, typerate_limit_error, messageShannon Coder call quota reached. Upgrade your plan at shannon-ai.com/plan. - On
/v1/responses,shannon-coder-1has no call allowance: it is charged in tokens from your balance at $8.00 per 1M, like every other model.
Flood protection
An account may send 120 requests per minute. That is the only limit on request rate, and it is the same on every plan. It exists to stop floods, not to slow normal use.
- The minute is a fixed window of 60 seconds that opens with your first request. When it ends, the count starts again at zero.
- The count is per account, not per key and not per IP address. Rotating the key does not open a new window.
- The 121st request inside a window is answered with status
429, typerate_limit_errorand the messageToo many requests. Retry in <N>s.Nis the number of seconds until the window ends, from 1 to 60. - Flood protection is checked before the balance. A request it refuses reserves nothing and costs nothing.
{
"error": {
"type": "rate_limit_error",
"message": "Too many requests. Retry in 37s."
}
} {
"type": "error",
"error": {
"type": "rate_limit_error",
"message": "Too many requests. Retry in 37s."
}
} | Request | Flood protection |
|---|---|
POST /v1/chat/completions, POST /v1/messages, POST /v1/responses | Counted, one per request. |
GET /v1/models, POST /v1/tokenize, POST /v1/messages/count_tokens | Not counted. |
shannon-coder-1 on /v1/chat/completions and /v1/messages | Counted by the Shannon Coder call allowance instead. |
A request answered with 401, or with 400 for an unknown model | Not counted. |
| A request refused by flood protection | Counted toward the window. Nothing is charged. |
Parallel requests
There is no limit on how many requests an account has open at the same time, and no error for sending requests in parallel. Requests that cannot start at once wait in line and are answered in turn.
- Each request counts toward the 120 per minute when it arrives, whether or not earlier requests have finished.
- Each request holds its own reservation until it ends. Twenty open requests with the default output budget hold 20 × 4,096 = 81,920 tokens of balance. If the reservations together are larger than your balance, the next request gets the
Quota exceededreply, even though the finished calls would have cost less. A smallermax_tokensholds less. - A request without streaming sends nothing until its answer is complete, so give your client a timeout that covers the wait. A stream keeps its connection open while it waits. Streaming
Limits of a single request
| Limit | Value | Applies to | At the limit |
|---|---|---|---|
| Request body | 32 MiB (33,554,432 bytes) | Every endpoint | Status 413, type invalid_request_error. |
Output budget: max_tokens, max_completion_tokens, max_output_tokens | 1 to 65,536. Default 4,096; for shannon-coder-1 on /v1/chat/completions and /v1/messages the default is 65,536. | Every model, as the amount reserved from your balance. As the limit on the length of the answer: the hosted open-weight models, shannon-1.6-lite, shannon-1.6-pro and shannon-coder-1. | A value outside the range is moved to the nearest end of the range. No error. |
Stop sequences: stop, stop_sequences | 4 strings | Hosted open-weight models | The first 4 non-empty strings are used. |
| Image or file given as a URL | 8 MiB, read within 20 seconds, at most 5 redirects, a public http or https address | Every endpoint that takes images or files | The request is answered without that part. No error. |
| Image or file sent inline (base64) | No limit of its own. It counts toward the 32 MiB request body. | Every endpoint that takes images or files | Status 413 for the whole request. |
text of POST /v1/tokenize | 4,000,000 bytes | /v1/tokenize | Status 413, type invalid_request_error, message text too long. |
messages of POST /v1/tokenize and the body of POST /v1/messages/count_tokens | The 32 MiB request body | Both counting endpoints | Status 413. |
| Context window | Per model: context_window in GET /v1/models | Every model | What happens to a longer conversation depends on the model. Models & pricing |
Web searches (web_search: true) | Per plan and day: Free 3, Plus 30, Standard 50, Pro 60. One search is counted for a request whose search found results. | Requests that set web_search: true | With none left the request is answered without search. No error. Built-in Web Search |
Errors
The replies of this page. On /v1/messages the same error object is wrapped as {"type": "error", "error": {…}}.
| Status | Type | Message | When, and what to do |
|---|---|---|---|
429 | rate_limit_error | Quota exceeded. Upgrade your plan at shannon-ai.com/plan | The reservation of the request does not fit into your balance. Wait for 00:00 UTC, top up credit, change plan, or send a smaller max_tokens. |
429 | rate_limit_error | Too many requests. Retry in <N>s. | More than 120 requests in the current minute. Wait N seconds and send again. |
429 | rate_limit_error | Shannon Coder call quota reached. Upgrade your plan at shannon-ai.com/plan | The Shannon Coder calls of the current 4-hour window are used up. |
429 | rate_limit_error | Shannon routes are temporarily busy. Please retry. | The model cannot take the request at this moment. Send it again after a short pause. |
503 | api_error | Could not verify your quota right now. Please retry. | Your balance could not be read. Nothing is charged; send the request again. On /v1/responses with a Shannon model the status is 500. |
413 | invalid_request_error | The request body is larger than 32 MiB. On the OpenAI-format endpoints the error object carries code: "request_too_large". | |
413 | invalid_request_error | text too long | text of POST /v1/tokenize is longer than 4,000,000 bytes. |