Home → Help

Requests die at exactly 100 seconds behind Cloudflare

Cloudflare's free and Pro plans cap a proxied HTTP response at 100 seconds. If the model has not finished by then the connection is cut, regardless of what your own timeouts say. Streaming avoids it because bytes keep arriving; a single non-streamed long completion does not.

What you are seeing

Why it happens

The give-away is the precision. Network problems and origin timeouts scatter — they fail at 40 seconds one time and 180 the next. A proxy limit fires at the same number every time. If your failures cluster at 100-101 seconds, stop looking at your own stack.

It is measured to the first byte of the response body, which is why streaming behaves so differently. With stream: true the first token usually arrives in a second or two and the clock never gets close. Without it, the proxy waits for the whole completion — and a long reasoning chain or a 4K image can easily pass 100 seconds.

This also explains a confusing symptom: the request appears to succeed on the server. The upstream finished, the origin logged 200, the tokens were billed. Only the hop between proxy and client was cut. Checking your provider's usage log will show a completed, charged request that your client never received.

Confirm it is this

Time the failure. The exact number is the diagnosis:

time curl -s -o /dev/null -w '%{http_code} %{time_total}s\n' \
  'YOUR_BASE_URL/chat/completions' \
  -H 'Authorization: Bearer YOUR_KEY' \
  -H 'Content-Type: application/json' \
  -d '{"model":"YOUR_MODEL","messages":[{"role":"user","content":"Write a 3000 word essay."}]}'

Around 100s with a 524 confirms it. Run the same request with "stream": true — if that one completes, the limit is the only thing that was wrong.

How to fix it

  1. Turn on streamingThe single most effective fix, and usually a one-line change. stream: true in the request body; every major SDK supports it. First token arrives in seconds, so the 100-second window never opens.
  2. Use an endpoint that is not proxiedSome providers publish a separate API hostname that bypasses the CDN for exactly this reason. If yours does, point long-running calls at it and keep the website on the proxied name.
  3. Cap the work per requestSetting max_tokens to something the model can finish inside the window turns an invisible cut into a predictable stop. Splitting one long generation into several calls does the same and is easier to retry.
  4. Do not raise your client timeout and call it fixedThe connection is being cut upstream of you. A 300-second client timeout produces the same failure 200 seconds later — it only makes the symptom slower to reproduce.
APICLAN publishes https://api.apiclan.us/v1 for exactly this case — it is not behind the 100-second proxy limit, so non-streamed long completions finish. The main https://apiclan.us/v1 endpoint is fine for streaming and for anything that returns quickly.

Related

Unexpected token '<' when calling an OpenAI-compatible API401 invalid API key — when the key looks right but still fails

Last checked 2026-10-01. Written from problems diagnosed on a live OpenAI-compatible gateway, not collected from other sites.