Home → Help

A streaming response stops in the middle of a sentence

A healthy stream ends with an explicit terminator — a chunk carrying finish_reason, then data: [DONE]. If your last chunk has neither, the connection was cut and the answer you have is partial. Nothing in the HTTP status will tell you this, because the status was sent before the body began.

What you are seeing

Why it happens

Streaming responses commit to 200 OK as soon as the first byte leaves. Everything after that is body. If the connection dies at token 900 of 1200, the client has already seen a successful status line and will not raise unless it happens to be checking the stream's own end marker.

This is why the failure feels random. Short answers finish inside whatever window the weakest hop allows; long ones do not. The cut-off point moves because it depends on timing, not on content.

Proxy idle timeouts are the usual cause, and they interact badly with reasoning models: during a long reasoning phase no tokens are emitted at all, so an intermediary that measures silence rather than total duration can decide the connection is dead while the model is still thinking.

Confirm it is this

Read the raw stream instead of an SDK's parsed object, and look at the tail:

curl -N -s 'YOUR_BASE_URL/chat/completions' \
  -H 'Authorization: Bearer YOUR_KEY' \
  -H 'Content-Type: application/json' \
  -d '{"model":"YOUR_MODEL","stream":true,
       "messages":[{"role":"user","content":"Count slowly from 1 to 400."}]}' \
  | tail -5

A complete stream ends with a chunk containing "finish_reason":"stop" followed by data: [DONE]. If the last line is an ordinary content delta, the stream was cut.

How to fix it

  1. Assert on the terminatorTrack whether you saw [DONE] or a chunk with a non-null finish_reason. Without one of them, treat the result as failed rather than short. This is the single change that makes the problem visible.
  2. Keep the connection warm during long pausesIf an intermediary is closing on idle, heartbeat comments in the SSE stream prevent it. Providers that send them make reasoning models far more reliable over the same network path.
  3. Do not retry blindly from the startA severed stream has already billed the tokens it produced. Retrying the whole prompt doubles the cost. Where the answer is chunkable, resume; where it is not, at least log the partial so the spend is accounted for.
  4. Test with a deliberately slow, long generationShort prompts hide this completely. Ask for something that takes a minute to produce and you will know within one attempt whether your path is stable.
APICLAN passes the upstream stream through without buffering it, so finish_reason and [DONE] arrive exactly as the provider sent them. Setup for every client is in the quickstart.

Related

Unexpected token '<' when calling an OpenAI-compatible API401 invalid API key — when the key looks right but still fails

Last checked 2026-10-01. Written from problems diagnosed on a live OpenAI-compatible gateway, not collected from other sites.