Pricing (as of the September 21, 2026 launch)
Verified on the docs.x.ai API pricing page. All amounts in USD per 1M tokens. Standard pricing is unchanged from Grok 4.6. "Cached input" is the cheaper rate applied when prompt content is served from the cache (prompt caching explained).
| Tier | Input | Cached input | Output | Notes |
|---|---|---|---|---|
| Standard (short <200K) | $2.00 | $0.50 | $6.00 | Default rate, same as Grok 4.6 |
| Long context (over 200K) | $4.00 | $1.00 | $12.00 | Applies to the whole request |
Grok 4.7 Fast (Cursor and Grok Build only — the same Grok 4.7 served on faster infrastructure at twice the standard token rates)
| Prompt length | Input | Cached input | Output |
|---|---|---|---|
| Below 200K | $4.00 | $1.00 | $12.00 |
| Above 200K | $6.00 | $1.50 | $18.00 |
Thresholds follow the official wording (below 200k / above 200k). For the exact treatment of a prompt of exactly 200K tokens, see the docs.x.ai pricing table.
Primary sources: docs.x.ai/developers/pricing (Last updated: September 21, 2026) · docs.x.ai Grok 4.7 · x.ai/news/grok-4-7 (checked 2026-09-23)
Worked cost examples (standard tier, USD)
Input rate x input tokens + output rate x output tokens. Once a request passes the 200K long-context threshold, the higher rates apply to the entire request.
| Case | Input tokens | Output tokens | Estimated cost (USD) |
|---|---|---|---|
| Short exchange | 100,000 | 10,000 | $0.26 |
| Same input, 80% cached | 100,000 (80,000 cache hits) | 10,000 | $0.14 |
| Exactly 200K (boundary) | 200,000 | 20,000 | $0.52 → $1.04 if long-context applies |
| Long context (above 200K) | 250,000 | 20,000 | $1.24 |
Excludes cache-write charges, server-side tools and the 1.1x US regional multiplier. For your own token counts, use the token calculator.
Model specifications
Sources: docs.x.ai Grok 4.7 · xAI announcement. xAI says the model “uses a new, larger base model compared to Grok 4.6” and is better at verifying its own work and managing longer context. Note: the x.ai and docs.x.ai sites are branded “SpaceXAI” as of 2026-09-23; because a corporate rename has not been confirmed first-hand, this page refers to the provider as xAI.
Performance — ahead of Grok 4.6 on all seven official benchmarks
From xAI’s published benchmark table, comparing Grok 4.7, Grok 4.6, GPT-5.6 Sol and Fable 5.1 (Anthropic). These are vendor-reported figures and test conditions are not necessarily identical.
| Benchmark (domain) | Grok 4.7 | Grok 4.6 | GPT-5.6 Sol | Fable 5.1 | Note |
|---|---|---|---|---|---|
| CursorBench 4.0 (SWE) | 46.3% | 40.4% | 41.7% | 51.8% | +5.9pt vs. previous gen |
| DeepSWE v1.1 (long-horizon SWE) | 71.0%* | 65.2% | 72.7% | 70.0% | 2nd of four (* high effort) |
| EEBench (electrical engineering) | 64.0% | 53.0% | 39.4% | 56.4% | Best of four |
| AA Briefcase v1.1 (multi-hour office work, Elo) | 1,657 | 1,546 | 1,487 | 1,678 | Fable 5.1 slightly ahead |
| Terminal-Bench 4.0 (terminal work) | 37.6% | 20.3% | 37.3% | 57.9% | ~1.9x vs. previous gen |
| Harvey Legal Agent Benchmark (legal) | 19.6% | 15.8% | 2.5% | 6.7% | Best of four (~7.8x GPT-5.6 Sol) |
| HealthBench Professional (clinical reasoning) | 56.7% | 48.5% | 60.5% | 62.1% | 3rd of four |
“Winner” marks the top score in each benchmark. DeepSWE v1.1 is Datacurve’s long-horizon software engineering benchmark (113 tasks across 5 languages).
- GDPval (Elo): Grok 4.7 (xhigh) 1,695 — vs. Fable 5.1 (max) 1,735, Grok 4.6 (high) 1,605 and GPT-6 Astra (max) 1,542. That is +90 over the previous generation, with xAI highlighting better documents and presentations.
- Pricing unchanged: $2 input / $6 output (xHigh), the same as Grok 4.6 (High) in the vendor’s own chart. The same chart lists GPT-5.6 Sol Max at $4/$20 and Fable 5.1 Max at $10/$50.
- Bottom line: the vendor claim is “improved across the board at the same price.” But in the four-model field Fable 5.1 still leads three of the seven benchmarks and GPT-5.6 Sol one, so Grok 4.7 is not the outright leader.
Source: xAI: Introducing Grok 4.7 (benchmark table and GDPval chart)
Third-party view — affordable, but not at the frontier
Independent measurements are worth reading alongside the vendor’s numbers.
- Mid-pack on the composite intelligence index: Grok 4.7 scores 46 on Artificial Analysis’ Intelligence Index, while Claude Fable 5.1 and GPT-6 each score 53, according to reporting by the-decoder (2026-09-21). The reported gap widens further on agentic coding.
- Token consumption can outweigh the low unit price: VentureBeat notes (2026-09-21) that a model with cheap per-token pricing can still cost more per finished workload if it burns substantially more reasoning tokens. Compare measured token usage, not just list prices.
- Wide distribution: Grok 4.7 began rolling out in GitHub Copilot on 2026-09-21 (VS Code, JetBrains, Xcode and more). It is also available via the Grok API, Cursor, Grok Build, third-party coding harnesses, and model routers and cloud platforms.
Sources: the-decoder · VentureBeat · GitHub Changelog (all checked 2026-09-23). Third-party figures depend on each tester’s methodology.
How it compares
Based on xAI’s own comparison chart. Each model is shown at a different reasoning-effort setting, so this is not a like-for-like comparison.
| Model | Setting | Input | Output | Vs. Grok 4.7 |
|---|---|---|---|---|
| Grok 4.7 | xHigh | $2.00 | $6.00 | — (baseline) |
| Grok 4.6 | High | $2.00 | $6.00 | Same |
| GPT-5.6 Sol | Max | $4.00 | $20.00 | Grok cheaper (half input, 70% less output) |
| Claude Fable 5.1 | Max | $10.00 | $50.00 | Grok cheaper (1/5 input, ~1/8 output) |
The table above reflects xAI’s published chart, i.e. top reasoning-effort tiers; it may differ from the standard rates we verify independently on each vendor’s pricing page. See our pricing comparison for standard rates across all models.
Source: xAI: Introducing Grok 4.7 · standard vendor rates via our pricing page (as of 2026-09-23)
Where the benchmarks show strength
Coding, agentic work and specialist domains.
Coding & code review
46.3% on CursorBench 4.0 (+5.9pt over the previous generation). Available day one in Cursor and Grok Build.
Long terminal sessions
Terminal-Bench 4.0 jumps from 20.3% to 37.6% (~1.9x). Good for multi-step shell automation.
Legal & specialist domains
19.6% on the Harvey Legal Agent Benchmark — best of the four models compared.
Electrical engineering
64.0% on EEBench — best of the four models compared.
Documents & presentations
1,657 on AA Briefcase and 1,695 Elo on GDPval; xAI highlights improvement on professional deliverables.
Long-context workloads
500K context window — but rates double above 200K, so manage the length/cost trade-off.
When Grok 4.7 is a good fit — and when it isn't
- A strong price-per-token pick at the frontier tier: $2 input / $6 output is clearly cheaper than GPT-5.6 Sol or Fable 5.1 at their top effort settings. xAI claims “twice as fast, at half the price of comparable models.”
- Judge total cost by measured tokens: high reasoning effort (default high up to xhigh) inflates output tokens and narrows the price gap. Lower
reasoning_effortper task difficulty. - Mind the 200K wall: long-context calls bill at $4/$12 ($6/$1.50/$18 with Fast). Using cached input ($0.50) is the main lever for real savings.
- No Batch discount: the 20% Batch API discount covers only grok-4.3 and the grok-4.20 family, so grok-4.7 is not eligible.
- If you need the absolute best: on the Artificial Analysis Intelligence Index cited here, Claude Fable 5.1 and GPT-6 (53 each) score ahead of Grok 4.7 (46); consider routing the hardest work to those.
Estimate your actual cost
Calculate token costs for Grok 4.7 and every other model at once — just enter your input and output token counts.
🧮 Open the token calculator →📊 All-model pricing · 🤖 Compare with Grok 4.6 · 🗓️ Release timeline · 🔗 xAI official pricing
⚠️ Disclaimer
- Pricing was verified on the official docs.x.ai pages on 2026-09-23 and may change without notice. Always confirm on the official pages before you commit.
- This site is for information only and is not a recommendation or agent for any provider.
- Benchmark figures are vendor-reported and do not guarantee real-world results; third-party figures depend on each tester’s methodology.