Home → Help
error code: 524Requests consistently fail at 100-101 seconds, never at 90 or 120The same request succeeds when sent directly to the origin IPA streaming response stops mid-sentence with no error in the server logread ECONNRESET / IncompleteRead after a long pauseThe give-away is the precision. Network problems and origin timeouts scatter — they fail at 40 seconds one time and 180 the next. A proxy limit fires at the same number every time. If your failures cluster at 100-101 seconds, stop looking at your own stack.
It is measured to the first byte of the response body, which is why streaming behaves so differently. With stream: true the first token usually arrives in a second or two and the clock never gets close. Without it, the proxy waits for the whole completion — and a long reasoning chain or a 4K image can easily pass 100 seconds.
This also explains a confusing symptom: the request appears to succeed on the server. The upstream finished, the origin logged 200, the tokens were billed. Only the hop between proxy and client was cut. Checking your provider's usage log will show a completed, charged request that your client never received.
Time the failure. The exact number is the diagnosis:
time curl -s -o /dev/null -w '%{http_code} %{time_total}s\n' \
'YOUR_BASE_URL/chat/completions' \
-H 'Authorization: Bearer YOUR_KEY' \
-H 'Content-Type: application/json' \
-d '{"model":"YOUR_MODEL","messages":[{"role":"user","content":"Write a 3000 word essay."}]}'
Around 100s with a 524 confirms it. Run the same request with "stream": true — if that one completes, the limit is the only thing that was wrong.
stream: true in the request body; every major SDK supports it. First token arrives in seconds, so the 100-second window never opens.max_tokens to something the model can finish inside the window turns an invisible cut into a predictable stop. Splitting one long generation into several calls does the same and is easier to retry.https://api.apiclan.us/v1 for exactly this case — it is not behind the 100-second proxy limit, so non-streamed long completions finish. The main https://apiclan.us/v1 endpoint is fine for streaming and for anything that returns quickly.Last checked 2026-10-01. Written from problems diagnosed on a live OpenAI-compatible gateway, not collected from other sites.