Home → Help

Token usage missing from a streamed response

Streaming omits usage by default. Set stream_options: {"include_usage": true} and the counts arrive in one extra chunk at the very end, after the last content delta.

What you are seeing

Why it happens

Usage cannot be known until generation stops, so there is nothing to put in the early chunks. Rather than send a placeholder, the API leaves the field null and adds a final chunk when the numbers exist.

That final chunk has an empty choices array. Code that stops reading as soon as it sees finish_reason, or that skips chunks with no delta, throws away the very chunk it needs.

Not every OpenAI-compatible implementation supports the option. Where it is missing you either estimate locally or read the authoritative numbers from the provider's own log.

Confirm it is this

Ask for usage and look at the last line of the stream:

curl -N 'YOUR_BASE_URL/chat/completions' \
  -H 'Authorization: Bearer YOUR_KEY' -H 'Content-Type: application/json' \
  -d '{"model":"MODEL","stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hi"}]}' | tail -3

The chunk before [DONE] should carry a populated usage object with an empty choices array.

How to fix it

  1. Request the option explicitlyAdd stream_options with include_usage set to true. Without it the field stays null no matter how you read the stream.
  2. Keep reading until [DONE]Do not break out of the loop on finish_reason. The usage chunk comes after it.
  3. Treat local counts as estimatesClient-side tokenisers drift from what the provider charges, especially around cached input. Reconcile against the provider's log, not your own counter.
APICLAN records every call server-side regardless of streaming — timestamp, model, input, output, cache read, cache write and the exact amount deducted. That log is the authoritative number for reconciliation.

Related

Unexpected token '<' when calling an OpenAI-compatible API401 invalid API key — when the key looks right but still fails

Last checked 2026-10-01. Written from problems diagnosed on a live OpenAI-compatible gateway, not collected from other sites.