Major LLM API Pricing Comparison [September 2026]
We reference each provider's official pricing page and list only re-verified figures. All prices in USD / 1M tokens. Use English calculator to estimate costs. See GPT-5.6 Luna deep dive →, or the glossary entry on prompt caching (cache hits).
💡 Understanding the terms
- Input vs. output tokens — tokens you send (prompt, system instructions, history) vs. tokens the model generates. Output is usually the more expensive meter, but rates vary by model.
- Cache hit (cached input) — reusable prompt content served from the provider's prompt cache at a reduced rate. Behavior varies by provider; "hit" generally means the content was already cached.
- Cache miss (fresh input) — content not present in the cache, billed at the standard input rate.
- Cache write — the charge applied when content is stored in the prompt cache. It may also recur when the cache expires and is refreshed. Cache TTLs, write rates, and refresh rules vary by provider (e.g., Anthropic: 5-minute and 1-hour cache durations).
- MoE (Mixture of Experts) — an architecture that routes each token to a subset of specialized "experts" instead of activating all parameters. MoE does not by itself guarantee a lower bill; pricing follows the provider's policy.
- Context window — the maximum tokens a model can process in a single request (input, and sometimes output). It varies by model; a larger context is not always better for cost.
- Long-context surcharge — an additional rate some providers apply once the billable context exceeds a threshold. Others use a flat rate across the full window; check each model's rules.
- Introductory / promo price — a temporary launch discount; the "list" (standard) price applies after the window closes. Promos may carry expiry dates or usage restrictions.
- Batch / Flex / Fast (Priority) — provider-specific processing modes, often cheaper (batch/flex, sometimes ~50% off) or faster (priority/fast). Availability and meaning vary by provider.
- Peak / off-peak — time-of-day pricing (e.g., DeepSeek) where off-peak hours may be half the peak rate. Timezone and weekday rules are set by the provider.
OpenAI
📋 Standard, short-context rate. Long-context: see official page
| Model | Input | Cached Input | Output | Notes |
|---|---|---|---|---|
| GPT-6 Astra New | $10.00 | $1.00 | $50.00 | New flagship (announced 2026-09-03). Long: $20/$2/$75. Fast mode 2× →Deep dive |
| GPT-6 Sol New | $2.00 | $0.20 | $10.00 | Released 2026-09-22. Successor to GPT-5.6 Sol at 50% off (was $4/$20). 1,050,000 context, 128K output →Deep dive |
| GPT-6 Luna New | $0.10 | $0.01 | $0.50 | Released 2026-09-22. Successor to GPT-5.6 Luna at 50% off (was $0.20/$1.20). Free/Go tiers get it in the desktop app →Deep dive |
| GPT-5.6 Sol Promo | $4.00 thru 11/21 | $0.40 | $20.00 | Previous generation of GPT-6 Sol (legacy). Promo (was $5/$30). Long: $8/$0.80/$30 →Deep dive |
| GPT-5.6 Terra ↓ | $2.00 | $0.20 | $12.00 | Cut 2026-07-30 (was $2.50/$15). Long: $4/$0.40/$18 |
| GPT-5.6 Luna ↓ | $0.20 | $0.02 | $1.20 | Previous generation of GPT-6 Luna (legacy). Cut 2026-07-30 (was $1/$6). Long: $0.40/$0.04/$1.80 →Deep dive |
| GPT-5.5 | $5.00 | $0.50 | $30.00 | Previous flagship |
| GPT-5.4 | $2.50 | $0.25 | $15.00 | Balanced (superseded by Terra) |
| GPT-5.4 mini | $0.75 | $0.075 | $4.50 | Lightweight tier |
| GPT-5.4 nano | $0.20 | $0.02 | $1.25 | Cheapest tier |
Primary source: developers.openai.com/api/docs/pricing (2026-08-29 JST). Priority processing renamed to Fast mode. 💰 GPT-5.6 Sol is at a promo price of $4/$20 (at least through 2026-11-21; was $5/$30). Cloudflare AI Gateway has a limited-time 50% off (through 2026-09-18, Unified Billing only; vs Cloudflare's listed $5/$30 standard). 📄 GPT-5.6 Sol: pricing, promo & competitor comparison → · 📄 GPT-5.6 Luna →
Anthropic (Claude)
📋 Standard API / text input / no cache
| Model | Input | Cache Write 5m | Cache Hit | Output | Notes |
|---|---|---|---|---|---|
| Claude Fable 5.1 New | $10.00 | $12.50 | $0.25 | $50.00 | Released 2026-09-01. Fable 5 successor. Cache read 75% off ($1.00→$0.25) →Deep dive |
| Claude Fable 5 | $10.00 | $12.50 | $1.00 | $50.00 | Previous-gen flagship. Cache read $1.00 |
| Claude Opus 5.5 New | $4.00 | $5.00 | $0.20 | $20.00 | Released 2026-09-22. First of the Claude 5.5 family. Cheaper than Opus 5 ($5/$25→$4/$20); cache reads down 60% ($0.50→$0.20). Fast mode $8/$40 →Deep dive |
| Claude Opus 5 New | $5.00 | $6.25 | $0.50 | $25.00 | Released 2026-07-24. Same price as Opus 4.8. Fast mode $10/$50. Previous generation to Opus 5.5 |
| Claude Opus 4.8 | $5.00 | $6.25 | $0.50 | $25.00 | Previous Opus. Fallback for Claude Max |
| Claude Sonnet 5.5 New | $2.00 | $2.50 | $0.20 | $10.00 | Released 2026-09-28. Same price as Sonnet 5 (generation update). Details → |
| Claude Sonnet 5 | $2.00 | $2.50 | $0.20 | $10.00 | Released 2026-06-30 (previous generation). $2/$10 now permanent (announced 2026-08-10) |
| Claude Sonnet 4.6 | $3.00 | $3.75 | $0.30 | $15.00 | Older balanced model |
| Claude Haiku 4.5 | $1.00 | $1.25 | $0.10 | $5.00 | High-volume, classification |
Primary source: platform.claude.com (Sonnet 5.5 rows verified against the official Sonnet 5.5 overview and pricing docs via Firecrawl on 2026-10-01; other Anthropic rows retrieved 2026-09-01 JST). Claude Sonnet 5 is now permanently $2/$10 (the $3/$15 increase scheduled for Sep 1 will not occur). Batch 50% off. Fable 5.1 cache read is $0.25 (75% below Fable 5's $1.00). Flat rate full 1M context. 📄 Claude Sonnet 5.5 deep dive → · 📄 Claude Fable 5.1 deep dive →
Google (Gemini Developer API)
📋 Paid tier / Standard / text, image, video input / standard API
| Model | Input | Input>200K | Output | Output>200K | Notes |
|---|---|---|---|---|---|
| Gemini 3.8 Flash New | $0.75 until 12/31 $1.50(1/1~) | — | $3.75 until 12/31 $7.50(1/1~) | — | Released 2026-09-02. Best reasoning/coding Flash. 3.8 Flash Cyber is Fairwind-only →Deep dive |
| Gemini 3.7 Flash New | $0.75 until 12/31 $1.50(1/1~) | — | $3.75 until 12/31 $7.50(1/1~) | — | Released 2026-08-14. Most capable Flash for coding/agents →Deep dive |
| Gemini 3.6 Flash ↓ | $0.75 until 12/31 $1.50(1/1~) | — | $3.75 until 12/31 $7.50(1/1~) | — | Cut to introductory price (was $1.50/$7.50) |
| Gemini 3.5 Flash | $1.50 | — | $9.00 | — | GA May 2026 |
| Gemini 3.5 Flash-Lite | $0.30 | — | $2.50 | — | GA July 2026. Cheapest Gemini 3.x |
| Gemini 3.1 Pro Preview Preview | $2.00 | $4.00 | $12.00 | $18.00 | Preview pricing |
| Gemini 3.1 Flash-Lite | $0.25 | — | $1.50 | — | Audio $0.50 |
| Gemini 2.5 Pro | $1.25 | $2.50 | $10.00 | $15.00 | |
| Gemini 2.5 Flash | $0.30 | — | $2.50 | — | Audio $1.00 |
| Gemini 2.5 Flash-Lite | $0.10 | — | $0.40 | — | Audio $0.30 |
Gemini 3.7/3.6 Flash introductory price ($0.75/$3.75) valid through Dec 31, 2026; standard price $1.50/$7.50 from Jan 1, 2027. Primary source: ai.google.dev (2026-08-14 JST)
DeepSeek
📋 DeepSeek-V4.1-Flash launched 2026-09-10 with across-the-board price cuts. V4 Pro continues after 2026-09-14 — DeepSeek withdrew its 9/14 routing notice
deepseek-flash: lower rates, native vision). V4 Flash and V4 Flash Vision Exp are retired into V4.1 Flash. deepseek-v4-pro, by contrast, keeps running after 2026-09-14 with billing unchanged — the 9/14 routing notice was withdrawn. Peak hours: 01:00-04:00 and 06:00-10:00 UTC, Monday-Friday (JST 10:00-13:00 / 15:00-19:00). Weekends are off-peak all day; off-peak is half the peak rate.| Model | Input (miss) | Input (hit) | Output | Notes |
|---|---|---|---|---|
| DeepSeek V4.1 Flash New (deepseek-flash) | $0.15 (off-peak) $0.30 (peak) | $0.003 (off-peak) $0.006 (peak) | $0.60 (off-peak) $1.20 (peak) | 1M ctx, 384K out. Native vision (image input). Released 2026-09-10 → guide |
| DeepSeek V4 Pro (-0813) · live | $0.66 (off-peak) $1.32 (peak) | $0.022 (off-peak) $0.044 (peak) | $1.98 (off-peak) $3.96 (peak) | These rates are current. DeepSeek states V4 Pro keeps running after 2026-09-14 with billing unchanged (the 9/14 routing notice was withdrawn). No vision; concurrency 500. → continuation details |
⏰ Peak hours: 01:00-04:00 and 06:00-10:00 UTC, Monday-Friday only (JST 10:00-13:00 / 15:00-19:00; Beijing 09:00-12:00 / 14:00-18:00). Weekends are off-peak all day. Off-peak is half the peak rate. Effective 2026-08-16 16:00 UTC. Old promo rates (V4 Pro $0.435/$0.87, V4 Flash $0.14/$0.28) retired. The legacy model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are retired and served by V4.1 Flash (billed at the Flash price). DeepSeek has not announced a release date for V4.1 Pro. Source: api-docs.deepseek.com (2026-09-13) · Change Log (2026-09-10)
xAI (Grok)
📋 Standard API rate. Fast variant is 2× (see official docs for cache/other modes)
| Model | Input | Output | Notes |
|---|---|---|---|
| Grok 4.7 New | $2.00 | $6.00 | Released 2026-09-21. Fast variant 2× ($4/$12) · long context $4/$12 above 200K →Deep dive |
| Grok 4.6 | $2.00 | $6.00 | Released 2026-08-12. Fast variant 2× ($4/$12) →Deep dive |
Primary source: x.ai/news/grok-4-7 · docs.x.ai/developers/pricing · x.ai/news/grok-4-6 ("Pricing starts at $2 per million input tokens and $6 per million output tokens") · x.ai/pricing (consumer/enterprise plans; see docs.x.ai for full API details)
Zhipu (Z.ai / GLM)
📋 GLM-5.3-Flash at list price (intro 50% promo ended 9/9). GLM-5.3/5.2 at standard pricing
| Model | Input | Cached Input | Output | Notes |
|---|---|---|---|---|
| GLM-5.3-Flash New | $0.15 | $0.03 | $0.50 | 320B/18B MoE, 1M ctx, 128K out, MIT. Natively multimodal →Deep dive |
| GLM-5.3 | $1.40 | $0.26 | $4.40 | Reasoning model. Cache storage limited-time free |
| GLM-5.2 | $1.40 | $0.26 | $4.40 | MIT license. Cache storage limited-time free |
Primary source: docs.z.ai/guides/overview/pricing (2026-09-09) · GLM-5.3-Flash announcement. 📄 GLM 5.3 Flash deep dive →
Alibaba (Qwen)
📋 QwenCloud (international) standard rates. Flagship Qwen3.8-Max and low-cost Qwen3.8-Flash
| Model | Input | Output | Notes |
|---|---|---|---|
| Qwen3.8-Flash New | $0.15 | $0.47 | Released 2026-08-26. 125B/6B active MoE, 256K→1M context, multimodal. Qwen4 preview →Deep dive |
| Qwen3.8-Max | $2.00 | $6.00 | Flagship (2.4T MoE, 1M context). Implicit cache $0.25 |
Primary source: qwen.ai/blog (Qwen3.8-Flash-Next) (2026-09-01). Pricing $0.15/$0.47 per the official blog (the official X post lists $0.16/$0.47). 📄 Qwen3.8-Flash deep dive → · 🇨🇳 China AI overview →
DeepSeek V4.1 Flash — steep price cuts and native vision; V4 Pro continues (9/14 notice withdrawn)
V4.1 Flash costs far less than V4 Pro and adds native image input. It was initially announced that from 2026-09-14 every request — including those to deepseek-v4-pro — would route to V4.1 Flash, but DeepSeek withdrew that notice and confirmed V4 Pro continues. New rates: $0.15/$0.60 off-peak, $0.30/$1.20 peak (cache-miss input / output) — about 77% lower on input, 70% lower on output and 86% lower on cached input versus V4 Pro. Off-peak and weekend runs cut cost further. See the V4.1 Flash guide, V4 Pro continuation and the calculator.
⚠️ Disclaimer
- Pricing verified 2026-08-29 (OpenAI), 2026-09-01 (Anthropic), 2026-09-09 (Zhipu/Z.ai), 2026-09-13 (DeepSeek), 2026-08-14 (Gemini), 2026-08-12 (Grok 4.6) and 2026-09-23 (Grok 4.7), others 2026-08-04.
- Pricing may change without notice.
- DeepSeek peak/off-peak pricing takes effect 2026-08-16 16:00 UTC. Weekends are off-peak all day. DeepSeek-V4.1-Flash launched 2026-09-10 with across-the-board price cuts. Requests to V4 Flash / V4 Flash Vision Exp are served by V4.1 Flash at the V4.1 Flash price. V4 Pro, by contrast, continues after 2026-09-14 with billing unchanged (DeepSeek withdrew the 9/14 routing notice).
- This site is informational only.