Home → Help

Reading a 429 properly — retry-after and the limit headers

A 429 normally carries a Retry-After header, in seconds, and often a set of headers naming which limit was exhausted and when it resets. Waiting that long clears it; retrying immediately extends it.

What you are seeing

Why it happens

Providers meter several dimensions at once — requests per minute, tokens per minute, and sometimes concurrent requests. Hitting the token limit while well under the request limit is why 'only some requests' fail.

Token limits count what you send plus what you reserve. A few requests with a large max_tokens can exhaust a token-per-minute budget while the request count looks trivial.

Immediate retries are counted too. A tight retry loop turns a one-second pause into a sustained lockout, which is the usual reason a transient 429 becomes permanent.

Confirm it is this

Print the headers rather than the body when you get a 429:

curl -s -D- -o /dev/null -X POST 'YOUR_BASE_URL/chat/completions' \
  -H 'Authorization: Bearer YOUR_KEY' -H 'Content-Type: application/json' \
  -d '{"model":"MODEL","messages":[{"role":"user","content":"hi"}]}' \
  | grep -i 'retry-after\|ratelimit\|^HTTP'

Honour Retry-After when it is present. When it is not, exponential backoff with jitter starting around one second is the safe default.

How to fix it

  1. Honour Retry-After before your own backoffIf the header says 12 seconds, waiting 12 seconds works. Your one-second backoff does not, and it costs you the next attempt too.
  2. Add jitterWithout a random offset, every client in your fleet retries in lockstep and rebuilds the spike you are backing off from.
  3. Check whether it is really a balance problemSome gateways return 429 for an empty account. That one never clears on retry — read the error body before assuming it is rate limiting.
On APICLAN a 429 from an empty balance and a 429 from upstream congestion say different things in the body. The first needs a top-up and will never succeed on retry; the second usually clears within seconds.

Related

Unexpected token '<' when calling an OpenAI-compatible API401 invalid API key — when the key looks right but still fails

Last checked 2026-10-01. Written from problems diagnosed on a live OpenAI-compatible gateway, not collected from other sites.