OpenAI-Compatible API 429 Errors: Limits, Quota and Retry
Diagnose API 429 errors by their origin and error body, estimate safe throughput from token limits, and prevent retry storms across workers and SDKs.
Published 2026-09-29 · Updated 2026-09-29 · KeepRouter Editorial · 4 minute read

A 429 response means a limit was reached, but the status alone does not identify the exhausted resource. Read the error body and determine which service returned it. A temporary request-rate limit may justify a delayed retry; exhausted quota or account funding usually requires a configuration or billing correction.
For a team running scheduled jobs, the first fix is often controlling how quickly workers start requests. Increasing the number of workers can make throughput worse when they all retry against the same limit.
Identify the layer that rejected the request
The OpenAI error-code guide distinguishes rate-limit and quota-related failures. Compatible services may use different error codes or statuses. Apply the destination's documented contract rather than treating every OpenAI-style JSON object as an OpenAI account error.
| Observation | Check first | Useful response |
|---|---|---|
| Too many requests | Request rate and concurrency | Queue and pace calls |
| Token limit reached | Prompt sizes and output allowance | Reduce or schedule token demand |
| Quota or funding exhausted | Correct account and service balance | Fix the account condition |
| Works for one key but not another | Key scope and project limits | Compare configuration |
| Only one model fails | Model-specific allocation or route | Inspect that model's limits |
| Failure starts after retry rollout | All retry layers | Remove amplification |
Record timestamp, model ID, HTTP status, sanitized error type and any request identifier. Keep authentication headers and customer prompts out of the diagnostic record. If a gateway returned the error, do not assume you can inspect or alter its upstream provider account.
Calculate which limit binds
Suppose an illustrative account allows 120 requests per minute and 30,000 tokens per minute. If a job uses about 2,000 input tokens and 500 output tokens, token capacity permits roughly twelve such jobs per minute. The request-count limit is not the bottleneck.
requests_per_minute = 120
tokens_per_minute = 30_000
estimated_tokens_per_job = 2_500
headroom = 0.75
capacity = min(
requests_per_minute,
tokens_per_minute / estimated_tokens_per_job,
)
planned_jobs_per_minute = int(capacity * headroom)
print(planned_jobs_per_minute) # 9These are hypothetical inputs, not KeepRouter limits. Real services can count input, output reservations, concurrent requests or other resources differently. The calculation is a planning approximation; compare it with actual returned limits and observed request sizes.
Average tokens are also insufficient when a small number of very large requests arrive together. Consider a separate queue for long-document jobs so they do not consume the budget reserved for short interactive requests.
Use bounded backoff, not synchronized retries
The OpenAI rate-limit guide recommends exponential backoff with randomness and notes that unsuccessful requests can still count toward limits. A tight loop of immediate retries can therefore prolong the problem.
When the destination supplies a valid retry delay, respect it within your application's total deadline. Otherwise use bounded exponential backoff with jitter. Stop after a defined attempt count and return a clear retryable state to the job system. If the response says the account has no quota, delaying the same request does not restore quota.
Choose one layer to own retries. For example, two SDK retries produce up to three HTTP attempts. If the job runner repeats that operation three times, a single logical task can produce nine attempts. Add a workflow-level retry and the multiplication can grow again.
Separate concurrency from throughput
Concurrency is the number of requests in flight; throughput is how many complete per unit of time. Slow responses can keep a high number of requests open without increasing useful work. A concurrency semaphore helps contain that load, while a rate or token budget controls starts over time.
For an interactive app, also set a maximum queue wait. A request that waits two minutes before starting may violate the user's expectation even if it eventually succeeds. For offline jobs, a longer queue may be acceptable, but duplicate execution still needs protection.
Do not rotate through accounts or keys to evade a service's limits. Keys are useful for isolation and accounting, not for overriding an account policy. If the workload legitimately requires more capacity, request appropriate limits or design a documented alternative path.
Recover without repeating completed work
Give each job an application-owned ID and persist its state. If a timeout occurs after the model or a tool may have completed, check whether the result already exists before starting over. This matters especially when the model triggers a downstream side effect.
Keep a failed job distinguishable from a completed job with an empty answer. Store attempt count and the final reason. A dashboard that shows only eventual success can hide repeated charges and long delays.
For KeepRouter errors, start with the error reference and service status. Use the cost calculator to include retry attempts in the budget. Switching models or gateways is useful only when the alternative supports the task and the failure policy is explicit; it is not a substitute for understanding why the original request was rejected.
Frequently asked questions
Will exponential backoff fix exhausted quota?
No. Backoff addresses transient contention. Exhausted quota or funding requires fixing the account condition described by the destination service.
Why am I limited below the requests-per-minute number?
Token budgets, concurrency or model-specific limits may bind earlier. Compare actual prompt sizes and output allowances with the service's documented counting rules.
Should every layer retry independently?
No. Decide which layer owns retries and set a total deadline and attempt budget. Independent nested retries can multiply requests for one task.
Sources reviewed
Article last reviewed 2026-09-29