Home → Help
content is "" but the request returned 200finish_reason: "length" on a short promptUsage shows output tokens spent with nothing to show for themThe same prompt works on a non-reasoning modelRaising max_tokens suddenly makes it workReasoning models emit two kinds of output tokens: the internal chain and the answer you see. Billing counts both. max_tokens limits both. A value that was generous for a non-reasoning model can be entirely consumed before the first visible character.
The failure is quiet by design. The API did what it was told: it generated until it hit the limit, then stopped. finish_reason: "length" is the only clue, and it is easy to miss when your code reads choices[0].message.content and finds a harmless empty string.
We hit this on an internal document-scoring tool. max_tokens was set to 2500 — fine for the previous model, nowhere near enough once the model started reasoning first. The symptom looked like a broken proxy for a while, because the requests succeeded, the latency was normal, and the cost was real. The bill was the thing that gave it away: charges with no text to match them.
Look at finish_reason and the token counts, not just the content:
curl -s 'YOUR_BASE_URL/chat/completions' \
-H 'Authorization: Bearer YOUR_KEY' \
-H 'Content-Type: application/json' \
-d '{"model":"YOUR_MODEL","max_tokens":64,
"messages":[{"role":"user","content":"What is 17 * 23? Answer with the number only."}]}' \
| python3 -m json.tool
"finish_reason": "length" together with an empty content and a non-zero output token count is the signature. Run it again with max_tokens at 4000 and the same prompt will answer.
Last checked 2026-10-01. Written from problems diagnosed on a live OpenAI-compatible gateway, not collected from other sites.