Home → Guides → Grok 4.7 API: price, context window and switching from 4.6
xAI released grok-4.7 in September 2026. The pricing story is short — nothing changed per token — but the 200k long-context tier and the cache rate are where the actual bill is decided, and both are easy to miss.
List prices from xAI's model documentation, per million tokens:
| Prompt size | Input | Cached input | Output |
|---|---|---|---|
| under 200k tokens | $2.00 | $0.50 | $6.00 |
| 200k tokens or more | $4.00 | $1.00 | $12.00 |
These are identical to grok-4.6. Moving from one to the other does not change
the per-token cost in either tier. The context window is 500k tokens.
What you pay here, drawn live from the current rates:
| Model | Input / 1M | Output / 1M |
|---|---|---|
grok-4.7 | $0.400 | $1.20 |
grok-4.6 | $0.400 | $1.20 |
grok-4.5 | $0.400 | $1.20 |
grok-4.3 | $0.250 | $0.500 |
xAI prices a request by the size of its prompt. Stay under 200,000 prompt tokens and the lower tier applies; reach 200,000 and the request is billed at the higher tier — input, cached input and output all double.
A worked example at list price. A prompt of 250,000 tokens with a 2,000-token answer:
Cutting 22% of the prompt cut the cost by 61%, because the whole request dropped into the cheaper tier. If a workload hovers around 200k — long documents, a large repository in context, an agent that keeps appending history — trimming to just under the line is the biggest single saving available. A 500k window is useful headroom; it is not a reason to fill it.
Cached input is a quarter of the input price in both tiers. Workloads that resend the same prefix — a fixed system prompt, a long document asked several questions, an agent loop over the same codebase — pay the cached rate for the repeated part.
The practical rule is to keep the front of the prompt stable. Anything that changes on every call near the top of the prompt (a timestamp, a shuffled tool list) stops the rest from being reused. Put volatile content at the end.
It is one string. The request shape, the endpoint and the key stay the same:
from openai import OpenAI
client = OpenAI(api_key="YOUR_KEY", base_url="https://apiclan.us/v1")
resp = client.chat.completions.create(
model="grok-4.7", # was "grok-4.6"
messages=[{"role": "user", "content": "Hello"}],
)
Grok is served through the OpenAI-compatible endpoint, so the base URL ends in
/v1. That also means Cursor, Cline, Roo Code, Cherry Studio and anything else
that takes an OpenAI base URL will call it without extra setup.
Before switching production traffic, confirm the model is reachable with your key — a key belongs to one group and only sees that group's models:
curl -s https://apiclan.us/v1/models -H "Authorization: Bearer YOUR_KEY" | grep grok
If grok-4.7 is not listed, the problem is routing rather than your request;
here is why that happens.
We sent a single deliberately trivial request through this platform on 25 September 2026 and kept the whole response. It is a more useful number than any benchmark, because it shows the floor — what you pay when you ask for almost nothing.
curl https://apiclan.us/v1/chat/completions -H "Authorization: Bearer YOUR_KEY" -H "Content-Type: application/json" -d '{"model":"grok-4.7",
"messages":[{"role":"user","content":"Reply with exactly: ok"}],
"max_tokens":16}'
The answer was the two characters we asked for. The accounting was not two characters:
| Field | Value |
|---|---|
model in the response | grok-4.7-build |
prompt_tokens | 1,247 |
of which cached_tokens | 1,152 |
completion_tokens | 41 |
of which reasoning_tokens | 40 |
| Total round trip | 1.54 s |
There is a floor, and it is about 1,300 tokens. A six-word prompt arrived at the model as 1,247 input tokens, because the model carries a large built-in prefix. If your plan is high-frequency classification or routing with one-line prompts, that floor — not your prompt — is what you are paying for. At list price this single call costs roughly $0.001. Ten thousand of them is $10 before you have sent anything substantial.
max_tokens did not cap the output. We asked for 16 and were billed for
41 completion tokens, 40 of which were reasoning. On this model the reasoning budget is not
constrained by that parameter, so sizing max_tokens down is not a way to make a
reasoning call cheap. If you need short and cheap, you need a non-reasoning model, not a
smaller limit.
Most of the prefix came back as a cache read. 1,152 of the 1,247 input tokens were billed at the cached rate rather than the input rate — a quarter of the price. That is why the floor is a tenth of a cent rather than half a cent, and it happens on a cold call, with no prompt caching set up on your side.
One caveat we will state plainly: this is one measurement, not a benchmark. It tells you the shape of the bill for a minimal request, and it is reproducible with the curl above. It does not tell you how the model behaves on your workload.
On cost there is nothing to decide: the rates are the same. That makes this a pure quality question, and the honest way to answer it is on your own inputs. We are not going to quote benchmarks here. Run twenty real prompts from your product through both model names, read the outputs side by side, and keep whichever serves you better — because the price is identical, it is a free experiment.
Per-model rates in every currency view and 14 languages are on the grok-4.7 pricing page.
No subscription, no monthly minimum, no sales call. Top up with USDT and spend what you use — 1 USDT gives you 2 credits of API balance.
Read the 30-second quickstartPrices quoted on this page are regenerated automatically from live billing data. Third-party terms are quoted from that party's own published documentation.