Home → Help
Cost per request varies wildly for similar promptsMedian and mean cost per call differ by several timesA short conversation costs more than a long oneAgent frameworks cost far more than the same prompts by handThe math using the headline per-token price does not reproduce the invoiceHeadline pricing is quoted per million tokens for input and output, which invites a single average. Real invoices are the sum of four different rates: input, output, cache read and cache write. When the mix shifts, the bill shifts without the token total changing much.
Cache direction is the part that surprises people. Reading from cache is cheap — it is the reason caching exists. Writing to it costs more than sending the tokens fresh. So a loop that alters the beginning of its context on every turn invalidates the cache and pays the premium every time, while a loop that appends to a stable prefix pays the discount.
Numbers from our own billing records make the scale concrete: across a sample of cached requests, ordinary input and output together were about 3% of all tokens processed, while cache writes and reads made up the rest. In one session of thirty requests, cache writes alone accounted for 63% of the charge — tokens that produced no output at all.
Pull one real request from your provider's usage log and split the four counters:
# Any provider that reports usage will expose these four fields.
# What to compare:
# input_tokens charged at the input rate
# output_tokens usually ~5x the input rate
# cache_read_tokens a fraction of the input rate
# cache_creation_tokens charged ABOVE the input rate
#
# If cache_creation dominates, the cost is context churn, not generation.
Work out what share of the charge each counter accounts for. If cache writes lead, the lever is prompt stability, not a cheaper model.
Last checked 2026-10-01. Written from problems diagnosed on a live OpenAI-compatible gateway, not collected from other sites.