Home → Help

context_length_exceeded — what actually counts toward the limit

Input plus requested output has to fit the window together. Your entire message history counts on every call, and max_tokens is reserved up front — so a request can fail on the reservation even when the prompt alone would fit.

What you are seeing

Why it happens

Chat APIs are stateless. Every turn resends the whole history, so a conversation that has been running for an hour is not sending one message, it is sending all of them. That is why code which worked in the morning starts failing in the afternoon with nothing changed.

max_tokens is a reservation, not a cap on what you get. Asking for 8,000 output tokens removes 8,000 from the budget before the prompt is even measured.

System prompts, tool definitions and few-shot examples are usually invisible in the diff but very much present in the count. A large tool schema can be thousands of tokens on every single call.

How to fix it

  1. Lower max_tokens to what you actually needThis is the fastest fix and the one people skip. If answers are two paragraphs, reserving 8,000 tokens buys nothing and costs you the room.
  2. Trim history, not the current messageKeep the system prompt, the last few turns and a short summary of what came before. Dropping the oldest turns recovers far more than shortening what you are asking now.
  3. Count tool definitionsSerialise your tool schema and measure it. Trimming unused tools often frees more room than anything you do to the conversation.
Context windows differ by model, and each model's page on APICLAN lists the window alongside the price. Trimming history also cuts the bill directly, since input tokens are billed on every turn.

Related

Unexpected token '<' when calling an OpenAI-compatible API401 invalid API key — when the key looks right but still fails

Last checked 2026-10-01. Written from problems diagnosed on a live OpenAI-compatible gateway, not collected from other sites.