Home → Help
Charged twice for what looked like one request
Your client gave up and retried while the first request was still generating. The provider finished both, so both are billed — the timeout cancelled your wait, not the work.
What you are seeing
Two identical entries in the usage log seconds apartYou saw one error and were billed for two completionsCost per feature is roughly double what your own counter saysIt happens more on long or reasoning-heavy requests
Why it happens
An HTTP timeout is a local decision. Unless the connection is actually closed and the server chooses to abort on disconnect, generation continues to completion and is charged.
Automatic retries in SDKs are on by default and are invisible in your own logging. Code that appears to make one call can make three.
Reasoning models make this far more likely, because the time before the first token is long enough to cross a default timeout while everything is working normally.
How to fix it
- Set the timeout above the slowest response you expectMeasure the p99 of your own traffic and give it headroom. Most duplicate charges are a 60-second default meeting a 70-second answer.
- Turn off automatic retries for expensive callsSet max_retries to 0 on the client and retry deliberately, where you can log it and decide whether the work is worth repeating.
- Stream long requestsStreaming gives you a first token quickly, which keeps idle timeouts from firing and tells you the request is alive.
Every call on APICLAN is logged separately with its own timestamp and charge, so two entries a few seconds apart with the same model and similar token counts is the signature to look for when a bill seems doubled.
Related
Last checked 2026-10-01. Written from problems diagnosed on a live
OpenAI-compatible gateway, not collected from other sites.