Home → Help
The answer ends mid-word with no exceptionNo [DONE] marker in the raw streamWorks on short answers, fails on long oneschunked encoding ended prematurely / IncompleteReadRetrying produces a different cut-off point each timeStreaming responses commit to 200 OK as soon as the first byte leaves. Everything after that is body. If the connection dies at token 900 of 1200, the client has already seen a successful status line and will not raise unless it happens to be checking the stream's own end marker.
This is why the failure feels random. Short answers finish inside whatever window the weakest hop allows; long ones do not. The cut-off point moves because it depends on timing, not on content.
Proxy idle timeouts are the usual cause, and they interact badly with reasoning models: during a long reasoning phase no tokens are emitted at all, so an intermediary that measures silence rather than total duration can decide the connection is dead while the model is still thinking.
Read the raw stream instead of an SDK's parsed object, and look at the tail:
curl -N -s 'YOUR_BASE_URL/chat/completions' \
-H 'Authorization: Bearer YOUR_KEY' \
-H 'Content-Type: application/json' \
-d '{"model":"YOUR_MODEL","stream":true,
"messages":[{"role":"user","content":"Count slowly from 1 to 400."}]}' \
| tail -5
A complete stream ends with a chunk containing "finish_reason":"stop" followed by data: [DONE]. If the last line is an ordinary content delta, the stream was cut.
[DONE] or a chunk with a non-null finish_reason. Without one of them, treat the result as failed rather than short. This is the single change that makes the problem visible.finish_reason and [DONE] arrive exactly as the provider sent them. Setup for every client is in the quickstart.Last checked 2026-10-01. Written from problems diagnosed on a live OpenAI-compatible gateway, not collected from other sites.