Home → Guides → Grok 4.7 API: price, context window and switching from 4.6

Grok 4.7 API: price, context window and switching from 4.6

xAI released grok-4.7 in September 2026. The pricing story is short — nothing changed per token — but the 200k long-context tier and the cache rate are where the actual bill is decided, and both are easy to miss.

Price at a glance

List prices from xAI's model documentation, per million tokens:

Prompt sizeInputCached inputOutput
under 200k tokens$2.00$0.50$6.00
200k tokens or more$4.00$1.00$12.00

These are identical to grok-4.6. Moving from one to the other does not change the per-token cost in either tier. The context window is 500k tokens.

What you pay here, drawn live from the current rates:

ModelInput / 1MOutput / 1M
grok-4.7$0.400$1.20
grok-4.6$0.400$1.20
grok-4.5$0.400$1.20
grok-4.3$0.250$0.500

The 200k threshold is the number that matters

xAI prices a request by the size of its prompt. Stay under 200,000 prompt tokens and the lower tier applies; reach 200,000 and the request is billed at the higher tier — input, cached input and output all double.

A worked example at list price. A prompt of 250,000 tokens with a 2,000-token answer:

Cutting 22% of the prompt cut the cost by 61%, because the whole request dropped into the cheaper tier. If a workload hovers around 200k — long documents, a large repository in context, an agent that keeps appending history — trimming to just under the line is the biggest single saving available. A 500k window is useful headroom; it is not a reason to fill it.

Caching

Cached input is a quarter of the input price in both tiers. Workloads that resend the same prefix — a fixed system prompt, a long document asked several questions, an agent loop over the same codebase — pay the cached rate for the repeated part.

The practical rule is to keep the front of the prompt stable. Anything that changes on every call near the top of the prompt (a timestamp, a shuffled tool list) stops the rest from being reused. Put volatile content at the end.

Switching from grok-4.6

It is one string. The request shape, the endpoint and the key stay the same:

from openai import OpenAI

client = OpenAI(api_key="YOUR_KEY", base_url="https://apiclan.us/v1")

resp = client.chat.completions.create(
    model="grok-4.7",          # was "grok-4.6"
    messages=[{"role": "user", "content": "Hello"}],
)

Grok is served through the OpenAI-compatible endpoint, so the base URL ends in /v1. That also means Cursor, Cline, Roo Code, Cherry Studio and anything else that takes an OpenAI base URL will call it without extra setup.

Before switching production traffic, confirm the model is reachable with your key — a key belongs to one group and only sees that group's models:

curl -s https://apiclan.us/v1/models -H "Authorization: Bearer YOUR_KEY" | grep grok

If grok-4.7 is not listed, the problem is routing rather than your request; here is why that happens.

What one real call actually costs

We sent a single deliberately trivial request through this platform on 25 September 2026 and kept the whole response. It is a more useful number than any benchmark, because it shows the floor — what you pay when you ask for almost nothing.

curl https://apiclan.us/v1/chat/completions   -H "Authorization: Bearer YOUR_KEY" -H "Content-Type: application/json"   -d '{"model":"grok-4.7",
       "messages":[{"role":"user","content":"Reply with exactly: ok"}],
       "max_tokens":16}'

The answer was the two characters we asked for. The accounting was not two characters:

FieldValue
model in the responsegrok-4.7-build
prompt_tokens1,247
  of which cached_tokens1,152
completion_tokens41
  of which reasoning_tokens40
Total round trip1.54 s

Three things worth knowing before you budget

There is a floor, and it is about 1,300 tokens. A six-word prompt arrived at the model as 1,247 input tokens, because the model carries a large built-in prefix. If your plan is high-frequency classification or routing with one-line prompts, that floor — not your prompt — is what you are paying for. At list price this single call costs roughly $0.001. Ten thousand of them is $10 before you have sent anything substantial.

max_tokens did not cap the output. We asked for 16 and were billed for 41 completion tokens, 40 of which were reasoning. On this model the reasoning budget is not constrained by that parameter, so sizing max_tokens down is not a way to make a reasoning call cheap. If you need short and cheap, you need a non-reasoning model, not a smaller limit.

Most of the prefix came back as a cache read. 1,152 of the 1,247 input tokens were billed at the cached rate rather than the input rate — a quarter of the price. That is why the floor is a tenth of a cent rather than half a cent, and it happens on a cold call, with no prompt caching set up on your side.

One caveat we will state plainly: this is one measurement, not a benchmark. It tells you the shape of the bill for a minimal request, and it is reproducible with the curl above. It does not tell you how the model behaves on your workload.

Should you move?

On cost there is nothing to decide: the rates are the same. That makes this a pure quality question, and the honest way to answer it is on your own inputs. We are not going to quote benchmarks here. Run twenty real prompts from your product through both model names, read the outputs side by side, and keep whichever serves you better — because the price is identical, it is a free experiment.

Per-model rates in every currency view and 14 languages are on the grok-4.7 pricing page.

Start using it

No subscription, no monthly minimum, no sales call. Top up with USDT and spend what you use — 1 USDT gives you 2 credits of API balance.

Read the 30-second quickstart

Prices quoted on this page are regenerated automatically from live billing data. Third-party terms are quoted from that party's own published documentation.