Home → Help

The bill is higher than the token count seems to justify

Multiply by the right rate. Output is usually about five times input. Cache reads are a small fraction of the input rate; cache writes are charged above it. A workload that keeps rewriting its context pays the write rate over and over while producing very little visible output.

What you are seeing

Why it happens

Headline pricing is quoted per million tokens for input and output, which invites a single average. Real invoices are the sum of four different rates: input, output, cache read and cache write. When the mix shifts, the bill shifts without the token total changing much.

Cache direction is the part that surprises people. Reading from cache is cheap — it is the reason caching exists. Writing to it costs more than sending the tokens fresh. So a loop that alters the beginning of its context on every turn invalidates the cache and pays the premium every time, while a loop that appends to a stable prefix pays the discount.

Numbers from our own billing records make the scale concrete: across a sample of cached requests, ordinary input and output together were about 3% of all tokens processed, while cache writes and reads made up the rest. In one session of thirty requests, cache writes alone accounted for 63% of the charge — tokens that produced no output at all.

Confirm it is this

Pull one real request from your provider's usage log and split the four counters:

# Any provider that reports usage will expose these four fields.
# What to compare:
#   input_tokens           charged at the input rate
#   output_tokens          usually ~5x the input rate
#   cache_read_tokens      a fraction of the input rate
#   cache_creation_tokens  charged ABOVE the input rate
#
# If cache_creation dominates, the cost is context churn, not generation.

Work out what share of the charge each counter accounts for. If cache writes lead, the lever is prompt stability, not a cheaper model.

How to fix it

  1. Stabilise the front of the promptAnything that changes on every turn — a timestamp, a shuffled tool list, a counter — invalidates the cache from that point on. Move volatile content to the end and the prefix stays cacheable.
  2. Look at output length before switching modelsOutput costs several times input. A prompt change that shortens answers by a third often saves more than moving to a cheaper tier, and needs no re-validation.
  3. Route cheap steps to a cheap modelClassification, extraction and routing rarely need a flagship. Sending only the final reasoning step to the expensive model usually cuts blended cost by more than half.
  4. Compare per-million-output, not per-million-tokenA blended per-token figure hides the thing that actually drives the bill. Two models with similar averages can differ sharply once the input/output ratio of your workload is applied.
APICLAN itemises all four counters for every call, so a surprising invoice can be traced to a specific request rather than estimated. Per-model rates are on the pricing pages.

Related

Unexpected token '<' when calling an OpenAI-compatible API401 invalid API key — when the key looks right but still fails

Last checked 2026-10-01. Written from problems diagnosed on a live OpenAI-compatible gateway, not collected from other sites.