Home → Help

Claude Code: API Error 400 prompt is too long

Everything the session has seen is resent on every turn. A few large file reads or one very long command output can take the conversation past the model's window, and from that point nothing you type is small enough to fit.

What you are seeing

Why it happens

Agent sessions are cumulative by design — the model needs earlier context to stay coherent. The cost is that a single 40,000-token file read stays in the transcript for the rest of the session.

Tool output is the usual culprit rather than your messages. A full log dump, a large JSON response or an unfiltered directory listing can be larger than everything you have typed all day.

Compaction helps but is not instant. Once you are past the limit, the request that would trigger compaction may itself be rejected, which is why the session appears wedged.

How to fix it

  1. Start a fresh session for a new taskThe cheapest fix, and the one that also reduces your bill: input tokens are charged on every turn, so a bloated session costs more per message as well as failing.
  2. Read less, filter morePrefer grep and targeted line ranges over reading whole files, and pipe long command output through head or a filter before it enters the transcript.
  3. Check the base URL if it started suddenlyA sudden change with no growth can mean you switched to a model with a smaller window. Confirm which model the session is actually using.
On APICLAN the same rule applies to billing as to the limit: every turn resends the whole conversation, so input tokens accumulate. Cache reads soften the cost but not the window — the tokens still count toward it.

Related

Unexpected token '<' when calling an OpenAI-compatible API401 invalid API key — when the key looks right but still fails

Last checked 2026-10-01. Written from problems diagnosed on a live OpenAI-compatible gateway, not collected from other sites.