Join our community of builders on Discord!

Limits

Each API key is held to two limits, enforced across all of the API's servers. The server sets them, the same for every key, and each key has its own budget at them: two keys of one wallet do not share a rate. A key carries no limits of its own. GET /api/api-keys shows how often each limit has refused each key (limitHits).
LimitDefaultCountsRefused with
Requests per minute60Every /v1 call, per calendar minute.429 rate_limit_exceeded, retry-after until the next minute
Completions in flight4Calls at once. A call beyond it is refused, not queued.429 concurrency_limit_exceeded, retry-after: 5
The limits are the server's configuration and may change; each refusal names the one it hit. Operators set them with DEVELOPER_API_MAX_REQUESTS_PER_MIN and DEVELOPER_API_MAX_CONCURRENT_SESSIONS. Minute windows are calendar minutes, so a burst either side of a minute's turn can pass up to twice the rate. Two more caps apply per key and are described with the key: the lifetime spendCapWei (402 spend_cap_exceeded, see API Keys) and the on-chain allowance of your wallet's delegate (see Payment).

Calls with no API key

A completion paid per call with no Authorization header (see Payment) has no key to hold it to. Its payer, the address its payment debits, takes the key's place, with limits the server sets for every payer alike:
LimitDefaultCountsRefused with
Requests per minute30Every paid call, per calendar minute, once its payment checks out.429 rate_limit_exceeded, retry-after until the next minute
Completions in flight1Calls at once. A call beyond it is refused before any session opens for it, and its payment is not settled.429 concurrency_limit_exceeded, retry-after: 5
The concurrency limit is small on purpose: a payment is checked before the API opens a session for the call, but settled only after, so each call in flight can cost the API a session that the payment never pays for. Mint an API key to run more at once. Operators set these with DEVELOPER_API_PAYER_REQUESTS_PER_MIN and DEVELOPER_API_PAYER_CONCURRENT_REQUESTS.

Server-wide

CapRefused with
New sessions opened per minute, across all keys, to leave workers time to take them on. A call opens one only when none of its key's sessions for the model is free; calls served by open sessions keep going. Spare sessions the server opens ahead of demand count too.429 session_open_limit_exceeded, retry-after until the next minute
60 calls a minute with a bad key, or with a per-call payment refused on a call with no key, from one address. Past that, every /v1 call from the address is refused until the minute turns, valid keys included. A call with no key and no payment, answered 402 with the requirements, is not counted.429 rate_limit_exceeded, retry-after

Per call

BoundRefused with
A job carries at most what fits in its one blob: about 124 KiB of the conversation, the JSON of its messages. An operator may set less with DEVELOPER_API_MAX_PROMPT_BYTES; a blob costs the same full or nearly empty, so the default cuts only a conversation that could not be sent at all. Past the bound the oldest turns are left out, never the system messages or the last turn; lightchain.dropped_messages says how many messages were. They go several at a time, so that the next calls' jobs start with the same turns and the worker's prompt cache keeps matching them: a long conversation's job carries between about half the bound, less a turn, and all of it.Not refused
The system messages and the last turn, which every job carries, must fit in one blob: about 124 KiB once encrypted.400 context_length_exceeded
A job has 5 minutes by default. A worker that takes longer, or fails, is taken for dead; if nothing had streamed yet, the API retries the call once in a new session, which may draw another worker.502 job_failed
When a call needs a new session, a worker must take it on within the claim window.503 no_worker_available
The chain's blob base fee must be at or below what the server pays to submit a prompt.503 blob_fee_too_high

Spare sessions

An operator can set DEVELOPER_API_WARM_SESSIONS (default 0, none) to keep that many free sessions ready in each session pool (a key's, or a payer's, sessions for one model), within the pool's concurrency limit. A call that leaves fewer free opens the rest in the background, after it has its own session, so the next calls take a ready session instead of waiting for one to open. Only calls do this: a pool nobody calls opens nothing more, and the very first call to a pool still waits for its own session. A pool held to one call at a time, as payers with no key are by default, keeps no spares. Each spare is one session open the API pays for: two transactions from its delegate, a requestSessionFor and a provideSessionKeyFor, about 600,000 gas together (measured on a devnet; about 0.0006 LCAI at a 1 gwei gas price). The job fees stay the caller's. A spare is then replaced like any session, once used past an hour old, and leaves after 12 hours. With the default of 0 the API opens a session only for a call that finds none free. Spares share the server's per-minute open cap with the opens calls need, so keep the number small. A pool's warm open that fails (no worker took the session on in the claim window, or the open cap refused it) pauses that pool's spares for a minute.

Handling refusals

  • 429: wait retry-after seconds. OpenAI SDKs retry 429 and 5xx by themselves, up to their retry count.
  • 503 and 502: retry with backoff (blob_fee_too_high clears when the chain's blob fee falls). A call that fails gives its fee back to the key's spend cap. On chain, a job that was submitted before the failure is settled by the protocol's job-timeout rules, and a retry is a new job.
  • 402: see Payment; retrying without paying will not help.
All refusals keep OpenAI's error shape, with the reason in error.code. Errors lists every code.