Home → Help

The model returns an empty string but you are still billed

On a reasoning model, max_tokens caps reasoning tokens and visible output together. If the chain of thought uses the whole budget, you get a valid 200 with an empty content field — and you are billed for every reasoning token.

What you are seeing

Why it happens

Reasoning models emit two kinds of output tokens: the internal chain and the answer you see. Billing counts both. max_tokens limits both. A value that was generous for a non-reasoning model can be entirely consumed before the first visible character.

The failure is quiet by design. The API did what it was told: it generated until it hit the limit, then stopped. finish_reason: "length" is the only clue, and it is easy to miss when your code reads choices[0].message.content and finds a harmless empty string.

We hit this on an internal document-scoring tool. max_tokens was set to 2500 — fine for the previous model, nowhere near enough once the model started reasoning first. The symptom looked like a broken proxy for a while, because the requests succeeded, the latency was normal, and the cost was real. The bill was the thing that gave it away: charges with no text to match them.

Confirm it is this

Look at finish_reason and the token counts, not just the content:

curl -s 'YOUR_BASE_URL/chat/completions' \
  -H 'Authorization: Bearer YOUR_KEY' \
  -H 'Content-Type: application/json' \
  -d '{"model":"YOUR_MODEL","max_tokens":64,
       "messages":[{"role":"user","content":"What is 17 * 23? Answer with the number only."}]}' \
  | python3 -m json.tool

"finish_reason": "length" together with an empty content and a non-zero output token count is the signature. Run it again with max_tokens at 4000 and the same prompt will answer.

How to fix it

  1. Budget for the reasoning, not just the answerA reasoning model needs room for both. If you want a 200-token answer, leaving 2000 tokens of headroom is not excessive — the chain is often several times longer than the reply.
  2. Treat finish_reason: length as an error in your codeDo not let an empty string through silently. Branch on it and either retry with a larger budget or fail loudly. This one check turns a mystery into a log line.
  3. Check whether the model supports a separate reasoning budgetSome expose a dedicated parameter so the visible answer gets its own allowance. Where that exists, use it instead of guessing one combined number.
  4. Compare against a non-reasoning model before blaming the networkIf the identical request returns text from a non-reasoning model, the transport is fine and the budget is the problem.
Every request on APICLAN shows its input, output and reasoning token counts in the usage log, so an empty answer with a real bill is visible rather than mysterious. Model-by-model behaviour is documented on the pricing pages.

Related

Unexpected token '<' when calling an OpenAI-compatible API401 invalid API key — when the key looks right but still fails

Last checked 2026-10-01. Written from problems diagnosed on a live OpenAI-compatible gateway, not collected from other sites.