Errors
Every failure carries a machine-readable code. The code is the contract — switch on it, never on the message text, which is written for humans and may be reworded at any time.
Native endpoints return { error: { code, message, details? } }. The /v1 surface returns OpenAI’s error object instead, so an OpenAI SDK raises the exception type it already knows. The code field is the same string in both.
Rate limits, quota and balance
free_allowance_exhaustedHTTP 429terminalThe free tier's MONTHLY message allowance is spent — it renews on the 1st. details carries limit, used and the upgrade path. 429, never 402.
NOT retryable today — the free allowance is a monthly one that renews on the 1st. Upgrade the account to keep going before then.
quota_exceededHTTP 429retryableA plan limit was hit. details.reason is concurrency (too many generations in flight — retry when one finishes) or daily_output_tokens.
Read details.reason. concurrency means too many generations are in flight for this plan — retry when one finishes, and prefer a bounded worker pool over a retry storm. daily_output_tokens means the day's budget is spent; it resets, so schedule rather than spin.
insufficient_creditsHTTP 429terminalThe balance will not cover the estimated cost of this request. Surfaced on /v1 as OpenAI's insufficient_quota. 429, never 402.
Not retryable by the client. The balance will not cover the request's estimated cost — add credit or reduce max_tokens. On /v1 this arrives as OpenAI's insufficient_quota type, which most SDKs already surface as a billing error rather than a rate limit.
rate_limitedHTTP 429retryableToo many requests. Back off and retry.
Back off exponentially with jitter and retry. This is the only 429 that a plain retry loop is the right answer to.
401 · 403
unauthorizedHTTP 401terminalMissing, malformed, or revoked API key.
Check the key is present, un-revoked, and sent as Authorization: Bearer …. If you just rotated, deploy the new value — there is no grace period on the old one.
key_expiredHTTP 401terminalThe API key's expiry time has passed. Mint a new key (the old id stays revoked) and retry with it; the request itself was otherwise valid.
Do not retry — the request cannot succeed unchanged.
insufficient_scopeHTTP 403terminalA valid key was presented but its scopes do not include the one this operation requires (e.g. a models:read key calling chat completions).
Do not retry — the request cannot succeed unchanged.
email_unverifiedHTTP 403terminalThe owning account has not confirmed its email address. Open the sign-in link that was mailed to it; the same key then works unchanged.
The account, not the key, is the problem. Open the sign-in link that was mailed to it; the same key then works unchanged.
account_suspendedHTTP 403terminalThe account is suspended and cannot start new generations.
Stop retrying and contact support. No key on this account will generate.
account_restrictedHTTP 403terminalThe account is paused pending a trust & safety review.
Stop retrying. The account is paused pending a review; retries do not shorten it.
400 · 404 · 410 · 413 · 422
validation_errorHTTP 422terminalThe request body or query is malformed, or names an unsupported option.
Fix the request. Retrying an identical body will fail identically. details names the offending field where the API can identify one.
model_not_foundHTTP 404terminalNo published model has that id. Call GET /v1/models for the live list.
Call GET /v1/models and use an id from the list. Do not hardcode ids that you have not seen in that response.
not_foundHTTP 404terminalThe resource does not exist, or is not visible to this caller.
Treat as permanent for this identifier. Do not retry the same URL.
model_removedHTTP 410terminalThe model was withdrawn. A tombstone, so crawlers and clients stop retrying.
Permanent. Drop the id from your configuration — this is a tombstone specifically so clients and crawlers stop retrying.
payload_too_largeHTTP 413terminalThe request body exceeds the accepted size.
Shorten the prompt or split the work. Retrying the same body cannot succeed.
409 — the model is not servable right now
model_unavailableHTTP 409retryableThe model exists but has no servable version right now.
Retryable, but not instantly — back off in seconds, or fail over to another model id.
model_cold_startHTTP 409retryableThe model is asleep, and free accounts run only on models that are already awake. details.warm_models lists what is awake; details.estimated_wait_s is how long it takes to wake.
The model is asleep. details.warm_models lists what is awake right now and details.estimated_wait_s is how long it takes to wake: either wait that long, or reroute to one that is already awake.
5xx
internal_errorHTTP 500retryableSomething failed on our side. Safe to retry.
Retry with backoff. If it persists across models, it is ours, not yours.
service_disabledHTTP 503retryableAn operator has paused new generations platform-wide.
An operator has paused new generations platform-wide. Retry with a long backoff; check the status page rather than hammering.
- Read
code. If the table above says terminal, surface it to a human and stop — retries cost money and change nothing. - For retryable codes, back off exponentially with jitter. Cap total attempts; a stuck queue is more visible than an infinite one.
- Bound concurrency yourself.
quota_exceededwithdetails.reason = "concurrency"is your worker pool being wider than your plan, and no retry schedule fixes that. - Log the
codeand the request id, not the prompt.
Each endpoint page lists only the codes that endpoint can return, with the exact response body — start at the endpoint index.