Home → Pricing → grok-4.6
xAI's current flagship. Notably cheap on output tokens, which makes it attractive for long-form generation.
| Per 1M tokens | List price | APICLAN | You save |
|---|---|---|---|
| Input | $2.00 | $0.550 | 72% |
| Output | $6.00 | $1.65 | 72% |
Cached input is billed at a fraction of the input rate, so agent tools that resend the same context — Codex, Claude Code, Cline — cost considerably less in practice than the raw token count suggests.
Put in the token counts from a typical request. Both columns update as you type.
A coding session through Codex on this platform used 90,459 tokens across five requests — roughly seventeen minutes of back-and-forth over a codebase. At list price that is $0.178. The customer paid $0.036.
That is real billing data, not an estimate. Heavy context reuse is exactly where the saving compounds: the same work, one fifth of the invoice.
The API is OpenAI-compatible. Point your base URL at APICLAN and pass the model name — nothing else in your code changes.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_KEY",
base_url="https://apiclan.us/v1",
)
resp = client.chat.completions.create(
model="grok-4.6",
messages=[{"role": "user", "content": "Hello"}],
)
Using Claude Code or the Anthropic SDK? Those clients append
/v1/messages themselves, so their base URL is
https://apiclan.us with no /v1. Full setup for every
client is in the Quickstart.
| Model | Input / 1M | Output / 1M |
|---|
Prices above are what you pay, not list price. Cheaper models in the same family often handle routine work at a fraction of the cost — worth testing before defaulting to the largest one.
No subscription, no monthly minimum, no sales call. Top up with USDT and spend what you use — 1 USDT gives you 2 credits of API balance.
Read the 30-second quickstartPricing verified 2026-08-24 · 按公开价目表. List prices change; this page is regenerated automatically from live billing data.