Get Started
Limits
Each API key is held to two limits, enforced across all of the API's servers. The server sets them, the same for every key, and each key has its own budget at them: two keys of one wallet do not share a rate. A key carries no limits of its own.GET /api/api-keys shows how often each limit has refused each key (limitHits).
The limits are the server's configuration and may change; each refusal names the one it hit. Operators set them with DEVELOPER_API_MAX_REQUESTS_PER_MIN and DEVELOPER_API_MAX_CONCURRENT_SESSIONS. Minute windows are calendar minutes, so a burst either side of a minute's turn can pass up to twice the rate.
Two more caps apply per key and are described with the key: the lifetime spendCapWei (402 spend_cap_exceeded, see API Keys) and the on-chain allowance of your wallet's delegate (see Payment).
Calls with no API key
A completion paid per call with noAuthorization header (see Payment) has no key to hold it to. Its payer, the address its payment debits, takes the key's place, with limits the server sets for every payer alike:
The concurrency limit is small on purpose: a payment is checked before the API opens a session for the call, but settled only after, so each call in flight can cost the API a session that the payment never pays for. Mint an API key to run more at once. Operators set these with DEVELOPER_API_PAYER_REQUESTS_PER_MIN and DEVELOPER_API_PAYER_CONCURRENT_REQUESTS.
Server-wide
Per call
Spare sessions
An operator can setDEVELOPER_API_WARM_SESSIONS (default 0, none) to keep that many free sessions ready in each session pool (a key's, or a payer's, sessions for one model), within the pool's concurrency limit. A call that leaves fewer free opens the rest in the background, after it has its own session, so the next calls take a ready session instead of waiting for one to open. Only calls do this: a pool nobody calls opens nothing more, and the very first call to a pool still waits for its own session. A pool held to one call at a time, as payers with no key are by default, keeps no spares.
Each spare is one session open the API pays for: two transactions from its delegate, a requestSessionFor and a provideSessionKeyFor, about 600,000 gas together (measured on a devnet; about 0.0006 LCAI at a 1 gwei gas price). The job fees stay the caller's. A spare is then replaced like any session, once used past an hour old, and leaves after 12 hours. With the default of 0 the API opens a session only for a call that finds none free. Spares share the server's per-minute open cap with the opens calls need, so keep the number small. A pool's warm open that fails (no worker took the session on in the claim window, or the open cap refused it) pauses that pool's spares for a minute.
Handling refusals
429: waitretry-afterseconds. OpenAI SDKs retry429and5xxby themselves, up to their retry count.503and502: retry with backoff (blob_fee_too_highclears when the chain's blob fee falls). A call that fails gives its fee back to the key's spend cap. On chain, a job that was submitted before the failure is settled by the protocol's job-timeout rules, and a retry is a new job.402: see Payment; retrying without paying will not help.
error.code. Errors lists every code.