Home → Guides → GPT-5.5 in production

GPT-5.5 in production

The most-called model on this platform — 3363 billed calls and counting. The interesting part is not the list price. It is that nine out of ten input tokens arrive as cache reads, which changes what it actually costs to run.

The short version

Why a $5 / $30 model can cost less than a $4 / $20 one

List prices compare tokens. Bills compare your tokens, and the two diverge as soon as caching enters. Cached input is billed at a fraction of the input rate, so a workload that re-reads the same prefix on every turn pays that fraction on most of what it sends.

In our own traffic gpt-5.5 reads 25,173,420 cached tokens against 2,807,198 fresh ones. Priced at the full input rate those cached tokens would cost many times what they actually did. That is the whole reason a model with a higher headline price shows a lower cost per call here than some of its cheaper-on-paper siblings.

The condition is stability. Caching only pays when the front of the prompt stays byte-identical between calls — same system prompt, same tool definitions, same document. Put a timestamp or a reshuffled tool list near the top and the discount disappears, because everything after the change has to be sent fresh again.

Latency: the reason it is still the default for interactive work

Measured on this platform, gpt-5.5 returns its first token in 3,035 and completes in 3.0 at the median. The reasoning-heavy models in the 5.6 family take substantially longer to start, which is fine for a batch job and very noticeable in a chat window.

If a human is waiting for the answer, time-to-first-token matters more than total tokens per second. That is the axis on which this model still wins inside our traffic, and it is why it remains the most-called one despite newer releases sitting next to it on the price list.

When to move to the 5.6 family instead

Three cases where the newer models are the right answer:

We are not publishing quality benchmarks here. We do not have our own evaluation data, and copying someone else's says nothing about your workload. Run twenty real prompts through both and read the outputs side by side — the price difference is small enough that the experiment is nearly free.

Start using it

No subscription, no monthly minimum, no sales call. Top up with USDT and spend what you use — 1 USDT gives you 2 credits of API balance.

Read the 30-second quickstart

Prices quoted on this page are regenerated automatically from live billing data. Third-party terms are quoted from that party's own published documentation.