Home → Help
429 Too Many Requests with no obvious patternRetries make the problem worseOnly some requests fail while others succeed at the same momentRate limiting under a load that used to be fineProviders meter several dimensions at once — requests per minute, tokens per minute, and sometimes concurrent requests. Hitting the token limit while well under the request limit is why 'only some requests' fail.
Token limits count what you send plus what you reserve. A few requests with a large max_tokens can exhaust a token-per-minute budget while the request count looks trivial.
Immediate retries are counted too. A tight retry loop turns a one-second pause into a sustained lockout, which is the usual reason a transient 429 becomes permanent.
Print the headers rather than the body when you get a 429:
curl -s -D- -o /dev/null -X POST 'YOUR_BASE_URL/chat/completions' \
-H 'Authorization: Bearer YOUR_KEY' -H 'Content-Type: application/json' \
-d '{"model":"MODEL","messages":[{"role":"user","content":"hi"}]}' \
| grep -i 'retry-after\|ratelimit\|^HTTP'
Honour Retry-After when it is present. When it is not, exponential backoff with jitter starting around one second is the safe default.
Last checked 2026-10-01. Written from problems diagnosed on a live OpenAI-compatible gateway, not collected from other sites.