Skip to content
Limits and balance

Limits and balance

Every request is served equally. No rate tiers. No separate API quota. You already paid for your tokens — use them as fast as you like.

This page explains what your balance is made of, what one request reserves and costs, how many requests you may send, and the few limits a single request can meet.

value of 1M tokens of balance
$5.00
daily allowance renews
00:00 UTC
flood protection, per account
120 requests / min

How requests are served

  • No rate tiers — One rule limits how fast requests may arrive, and it is the same for every account and every plan: 120 requests per minute. There is no limit on tokens per minute.
  • No separate API quota — The API spends the same balance as chat. A plan sets the size of today's allowance. It does not set a request rate.
  • As fast as you like — Requests sent in parallel are accepted and wait in line. They are not refused for being parallel.

Your balance

Your balance is counted in tokens. 1,000,000 tokens of balance are worth $5.00, and every price on the Models & pricing page is a rate against that value.

At any moment the balance is the sum of two parts.

  • Today's plan allowance — A number of tokens set by your plan. It is new every day at 00:00 UTC. What is left at the end of a day does not carry over.
  • Purchased credit — Tokens you bought as a pack. Credit does not expire and works on every plan, including Free.
Plan Tokens per day Worth
Free 30,000 $0.15
Plus 80,000 $0.40
Standard 265,000 $1.325
Pro 665,000 $3.325
  • Order of spending — Every request spends today's plan allowance first. Purchased credit is used only for what goes beyond the allowance on that day.
  • Chat and API share it — There is one balance per account. An API key spends from the balance of the account that owns it, at the same prices as chat.
  • Packs — Credit is sold in packs of 1,000,000 ($5.00), 2,000,000 ($10.00) and 5,000,000 ($25.00) tokens, or as an amount of your choice from 1,000,000 to 100,000,000 tokens at $5.00 per 1,000,000.

Top up credit Change plan

What a request reserves and what it costs

  • Reserve — When a request arrives it reserves its output budget from your balance: max_tokens on /v1/chat/completions and /v1/messages, max_output_tokens on /v1/responses. /v1/chat/completions also reads max_completion_tokens. The default is 4,096 and the range is 1 to 65,536.
  • Admit — The request is accepted only if the reservation fits into what is left of your balance. A balance that is above zero but smaller than the output budget gets the Quota exceeded reply. Send a smaller max_tokens to use the rest.
  • Settle — When the answer is complete, the reservation is replaced by the real charge. The charge can be lower or higher than the reservation.
  • Return — A request that ends with an error status gives its reservation back in full.

The real charge depends on the model family.

Models What is charged
Shannon models usage.total_tokens at the model's price per 1M. Input and output have one rate.
Hosted open-weight models Uncached input at the input rate, cached input at the cached rate, output at the output rate.

The amount in USD is taken from your balance in tokens at $5.00 per 1,000,000, rounded to a whole token.

Counting tokens with POST /v1/tokenize or POST /v1/messages/count_tokens is free and reserves nothing. Token counting

Where to see balance and usage

The Keys & usage page shows what you can spend right now, today's plan allowance, your purchased credit and the API spend of the last 30 days. Below that it lists every request your key made: time, endpoint, model, cached input, billed tokens and cost. Keys & usage

Every reply also carries a usage object with the token counts of that call.

Endpoint Fields of usage Added by hosted open-weight models
/v1/chat/completions prompt_tokens, completion_tokens, total_tokens prompt_tokens_details.cached_tokens, completion_tokens_details.reasoning_tokens
/v1/messages input_tokens, output_tokens cache_read_input_tokens, cache_creation_input_tokens
/v1/responses input_tokens, output_tokens, total_tokens input_tokens_details.cached_tokens, output_tokens_details.reasoning_tokens
  • usage holds the token counts of the model. The amount taken from your balance is not in the reply: it is the Billed tokens column of the request list on Keys & usage.
  • On /v1/messages with a hosted open-weight model, input_tokens is the uncached part of the input, cache_read_input_tokens is the cached part and cache_creation_input_tokens is always 0.
  • A stream on /v1/chat/completions carries usage in its last chunk before [DONE]. Streaming

When the balance runs out

A request whose reservation does not fit into your balance is answered with status 429, type rate_limit_error and the message below. Nothing is charged. The same reply is sent when the balance is above zero but smaller than the output budget of the request.

{
  "error": {
    "type": "rate_limit_error",
    "message": "Quota exceeded. Upgrade your plan at shannon-ai.com/plan"
  }
}

On /v1/responses the error object can also hold code and param, both null.

What you can do:

  • Wait for the next plan allowance at 00:00 UTC.
  • Top up credit. Credit is spent after the plan allowance and does not expire. Top up credit
  • Change to a plan with a larger daily allowance. Change plan
  • Send a smaller max_tokens, if some balance is left: the reservation is then smaller.

Shannon Coder call allowance

shannon-coder-1 on /v1/chat/completions and /v1/messages is counted in calls, not in tokens. Each plan includes a number of calls per 4-hour window. One request is one call.

Plan Calls per 4-hour window
Free 3
Plus 20
Standard 40
Pro 60
  • Windows start at 00:00, 04:00, 08:00, 12:00, 16:00, 20:00 UTC. Calls that are left at the end of a window do not carry over.
  • A call is counted when the request is accepted, before the model answers. A request that fails afterwards still counts as a call.
  • These calls reserve no tokens and take nothing from your balance. The request list on Keys & usage shows their token count and its value at the listed price.
  • The default max_tokens of shannon-coder-1 on these two endpoints is 65,536.
  • With no calls left the reply is status 429, type rate_limit_error, message Shannon Coder call quota reached. Upgrade your plan at shannon-ai.com/plan.
  • On /v1/responses, shannon-coder-1 has no call allowance: it is charged in tokens from your balance at $8.00 per 1M, like every other model.

Flood protection

An account may send 120 requests per minute. That is the only limit on request rate, and it is the same on every plan. It exists to stop floods, not to slow normal use.

  • The minute is a fixed window of 60 seconds that opens with your first request. When it ends, the count starts again at zero.
  • The count is per account, not per key and not per IP address. Rotating the key does not open a new window.
  • The 121st request inside a window is answered with status 429, type rate_limit_error and the message Too many requests. Retry in <N>s. N is the number of seconds until the window ends, from 1 to 60.
  • Flood protection is checked before the balance. A request it refuses reserves nothing and costs nothing.
{
  "error": {
    "type": "rate_limit_error",
    "message": "Too many requests. Retry in 37s."
  }
}
Request Flood protection
POST /v1/chat/completions, POST /v1/messages, POST /v1/responses Counted, one per request.
GET /v1/models, POST /v1/tokenize, POST /v1/messages/count_tokens Not counted.
shannon-coder-1 on /v1/chat/completions and /v1/messages Counted by the Shannon Coder call allowance instead.
A request answered with 401, or with 400 for an unknown model Not counted.
A request refused by flood protection Counted toward the window. Nothing is charged.

Parallel requests

There is no limit on how many requests an account has open at the same time, and no error for sending requests in parallel. Requests that cannot start at once wait in line and are answered in turn.

  • Each request counts toward the 120 per minute when it arrives, whether or not earlier requests have finished.
  • Each request holds its own reservation until it ends. Twenty open requests with the default output budget hold 20 × 4,096 = 81,920 tokens of balance. If the reservations together are larger than your balance, the next request gets the Quota exceeded reply, even though the finished calls would have cost less. A smaller max_tokens holds less.
  • A request without streaming sends nothing until its answer is complete, so give your client a timeout that covers the wait. A stream keeps its connection open while it waits. Streaming

Limits of a single request

Limit Value Applies to At the limit
Request body 32 MiB (33,554,432 bytes) Every endpoint Status 413, type invalid_request_error.
Output budget: max_tokens, max_completion_tokens, max_output_tokens 1 to 65,536. Default 4,096; for shannon-coder-1 on /v1/chat/completions and /v1/messages the default is 65,536. Every model, as the amount reserved from your balance. As the limit on the length of the answer: the hosted open-weight models, shannon-1.6-lite, shannon-1.6-pro and shannon-coder-1. A value outside the range is moved to the nearest end of the range. No error.
Stop sequences: stop, stop_sequences 4 strings Hosted open-weight models The first 4 non-empty strings are used.
Image or file given as a URL 8 MiB, read within 20 seconds, at most 5 redirects, a public http or https address Every endpoint that takes images or files The request is answered without that part. No error.
Image or file sent inline (base64) No limit of its own. It counts toward the 32 MiB request body. Every endpoint that takes images or files Status 413 for the whole request.
text of POST /v1/tokenize 4,000,000 bytes /v1/tokenize Status 413, type invalid_request_error, message text too long.
messages of POST /v1/tokenize and the body of POST /v1/messages/count_tokens The 32 MiB request body Both counting endpoints Status 413.
Context window Per model: context_window in GET /v1/models Every model What happens to a longer conversation depends on the model. Models & pricing
Web searches (web_search: true) Per plan and day: Free 3, Plus 30, Standard 50, Pro 60. One search is counted for a request whose search found results. Requests that set web_search: true With none left the request is answered without search. No error. Built-in Web Search

Errors

The replies of this page. On /v1/messages the same error object is wrapped as {"type": "error", "error": {…}}.

Status Type Message When, and what to do
429 rate_limit_error Quota exceeded. Upgrade your plan at shannon-ai.com/plan The reservation of the request does not fit into your balance. Wait for 00:00 UTC, top up credit, change plan, or send a smaller max_tokens.
429 rate_limit_error Too many requests. Retry in <N>s. More than 120 requests in the current minute. Wait N seconds and send again.
429 rate_limit_error Shannon Coder call quota reached. Upgrade your plan at shannon-ai.com/plan The Shannon Coder calls of the current 4-hour window are used up.
429 rate_limit_error Shannon routes are temporarily busy. Please retry. The model cannot take the request at this moment. Send it again after a short pause.
503 api_error Could not verify your quota right now. Please retry. Your balance could not be read. Nothing is charged; send the request again. On /v1/responses with a Shannon model the status is 500.
413 invalid_request_error The request body is larger than 32 MiB. On the OpenAI-format endpoints the error object carries code: "request_too_large".
413 invalid_request_error text too long text of POST /v1/tokenize is longer than 4,000,000 bytes.