DeepSeek R1 released — the origin of the "DeepSeek Shock"
Open-weight reasoning model released under the MIT license. Achieved OpenAI o1-class reasoning at low cost, triggering the "DeepSeek Shock."
Release timeline of major OpenAI (GPT-5), Anthropic (Claude 4/5), Google (Gemini 3), DeepSeek (R1/V4) and xAI (Grok) models, verified against each company's official announcements. See Pricing Comparison for current rates and the Token Calculator for cost estimates.
Open-weight reasoning model released under the MIT license. Achieved OpenAI o1-class reasoning at low cost, triggering the "DeepSeek Shock."
OpenAI's first unified system that auto-selects between a fast-response model and a reasoning model via a real-time router. Current generation is the GPT-5.6 family.
1M-token context window and enhanced multimodal understanding. Recorded 1501 Elo on LMArena.
Introduced a 1M-token context window to Anthropic's flagship. Long-document retrieval accuracy significantly improved.
Balanced model with upgraded coding, computer use, long-form reasoning and agentic planning. $3.00/$15.00
Major abstract reasoning gains. 77.1% on ARC-AGI-2. $2.00/$12.00 (≤200K)
Native Computer Use (OSWorld 75%). Five-level reasoning effort control. $2.50/$15.00
Coding (SWE-bench Pro 64.3%) and vision gains. Began adopting a new tokenizer.
Then-flagship with stronger agentic computer use and scientific research support. $5.00/$30.00
1M-token context, MIT license. V4 Pro (1.6T/49B active) · V4 Flash (284B/13B active). V4 Pro $0.435/$0.87 (promo at the time, now retired)
Optimized for agents and coding. Beats 3.1 Pro benchmarks at 4× speed. $1.50/$9.00
Code-defect misses reduced to roughly 1/4. Dynamic Workflows research preview. $5.00/$25.00
New tier above Opus. Fable 5 general availability, Mythos 5 limited. Paused 6/12 → restored 7/1. Fable 5 $10.00/$50.00
The "most agentic" Sonnet. Near-Opus 4.8 performance at a low price. The $2/$10 introductory price was made permanent on 2026-08-10 (the $3/$15 increase scheduled for Sep 1 will not occur). $2.00/$10.00 (permanent)
New flagship Sol, balanced Terra, and fastest low-cost Luna released together. Terra/Luna price cuts on 7/30. Sol $5/$30 · Terra $2/$12 · Luna $0.20/$1.20 (after 7/30)
3.6 Flash has cheaper output than 3.5 Flash ($7.50 vs $9.00). 3.5 Flash-Lite is the cheapest Gemini 3.x ($0.30/$2.50) for high-throughput agents. 3.5 Flash Cyber is cybersecurity-focused. 3.6 Flash $1.50/$7.50 · Flash-Lite $0.30/$2.50
Premium successor to Opus 4.8. Same price ($5/$25) with improved coding and professional tasks. Fast mode ($10/$50). ~85% fewer safety-classifier interventions vs Fable 5. $5.00/$25.00
Terra: $2.50→$2.00 / $15.00→$12.00. Luna: $1.00→$0.20 / $6.00→$1.20 (80% cut). Priority processing → Fast mode renamed the same day. Terra $2/$12 · Luna $0.20/$1.20
API model name deepseek-v4-flash. Native Responses API and Codex support. Agent benchmarks far exceed V4-Pro-Preview. V4-Pro API and APP/WEB unchanged.
Successor to Grok 4.5. Scores 61 on the AA Intelligence Index, matching GPT-5.6 Sol. 500K context, knowledge cutoff Feb 1 2026. $2/$6 (fast 2×)
GA rollout on APP, Web and API (model name deepseek-v4-pro). Responses API support, three thinking-effort levels (low/high/max). Peak/off-peak pricing from 8/16 16:00 UTC (off-peak = half of peak). Weekends off-peak all day since 8/23. V4 Pro output off-peak $1.98 / peak $3.96
Three weeks after 3.6 Flash. Intro price at half ($0.75/$3.75 thru 12/31 → $1.50/$7.50). FrontierCode 43.6% · DeepSWE 65.3%. $0.75/$3.75 (thru 12/31)
Routing to openai/gpt-5.6-sol through Cloudflare AI Gateway applies 50% off automatically (Unified Billing only; Bring Your Own Keys excluded). Input $2.50 / output $15 / cache read $0.25 (through 2026-09-18).
Experimental multimodal vision model, accessed via model='deepseek-v4-flash-vision-exp'. On par with V4 Flash on pure text; large gains on vision agent benchmarks.
Flagship Sol cut from $5→$4 input / $30→$20 output for 3 months. Applies to Fast mode, long-context, and Batch/Flex too. Was $5/$30. Promo guaranteed "at least through 2026-11-21".$4/$20 (Short context)
The first natively multimodal GLM-5 model. 320B/18B MoE, 1M context, 128K output, MIT. Beats GLM-5.2 on coding/agentic benchmarks at ~1/10 the price. $0.15/$0.50 (list)
A 125B (6B active) + 51B N-gram multimodal MoE. The GDN+QSA hybrid beats Qwen3.7-Plus on coding, office and agentic benchmarks at ~1/9 the training cost. Open-weight Flash-Next also on HF/ModelScope. $0.15/$0.47 (international)
A new frontier for coding, knowledge work and long-horizon agents. Same model, different safeguards (Fable 5.1 generally available / Mythos 5.1 trusted-access only). Cache reads cut 75% ($1.00→$0.25). 1M context, 128K output. Terminal-Bench-Science 52.6% (more than double Fable 5's 24.7%). $10/$50 · cache read $0.25
Three weeks after 3.7 Flash — third Flash release in six weeks. Intro $0.75/$3.75 (thru 12/31 → $1.50/$7.50). DeepSWE 73.8% · HLE-Verified 54.9%. Cybersecurity-focused 3.8 Flash Cyber is Fairwind Program-only. $0.75/$3.75 (thru 12/31)
Successor to GPT-5.6 Sol. State-of-the-art computer use, coding, cybersecurity and science. ARC-AGI-3 99.9% · ExploitBench 100% · Terminal-Bench Science 64.6%. 1.05M context, 128K output. Staged rollout (Daybreak → API/Plus). $10/$50
The smallest model in the new architecture family: 1M context, 384K output, native image input, thinking by default. API name deepseek-flash. V4 Flash and V4 Flash Vision Exp are retired and consolidated into V4.1 Flash. DeepSeek initially said V4 Pro requests would route to V4.1 Flash from 2026-09-14, but that notice was later withdrawn and V4 Pro continues (no V4.1 Pro release date announced). $0.15/$0.60 (off-peak, cache-miss input/output)
Footnote (2) of the official pricing page and the change log now read: "In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." The V4 Pro price row (off-peak input $0.66 / output $1.98) is retained, so deepseek-v4-pro keeps its own rates. $0.66/$1.98 (off-peak, cache-miss input/output)
Successor to Grok 4.6. 500K context, May 2026 knowledge cutoff. Beats Grok 4.6 on all seven official benchmarks and leads the four-model field on EEBench and the Harvey Legal Agent Benchmark. Third-party testing puts it at 46 on the Artificial Analysis Intelligence Index versus 53 for Fable 5.1 and GPT-6. $2/$6 (Fast 2x · $4/$12 above 200K)
Anthropic's first model in the new Claude 5.5 family. The official claim: performance at the level of Fable 5.1 on most work at 40% lower cost than Opus 5 on typical workloads. Input $5→$4, output $25→$20, cache reads $0.50→$0.20 (down 60%); Fast mode $8/$40. 1M context, 128K max output, June 2026 knowledge cutoff. Thinking can no longer be disabled — depth is set with effort (default medium). Sonnet 5.5 shipped on 9/28 (see the entry below); Haiku 5.5 is still unconfirmed. $4/$20 (cache read $0.20, write $5)
Three variants: Pro, Flash and Pro-UltraSpeed. Its official model page confirms text, image, video and audio input, text output and a 1M-token context. Pricing: Flash $0.14/$0.28, Pro $0.435/$0.87, Pro-UltraSpeed $4.35/$8.7 (cache hits $0.0028 / $0.0036 / $0.036). Xiaomi claims 46.32 on the Artificial Analysis Intelligence Index v4.3 and calls it “the strongest open-source model to date”. Technical report, training environment and RL code are released, with the production RL run live-streamed. $0.14 / $0.435 / $4.35 (input per 1M)
The two practical tiers of the GPT-6 generation. Sol falls from $4/$20 to $2/$10; Luna from $0.20/$1.20 to $0.10/$0.50. Cached input is discounted 90% ($0.20 for Sol, $0.01 for Luna). 1,050,000-token context, 128K max output. GPT-5.6 Sol and Luna become legacy. Available in ChatGPT Work and Codex; Free and Go users get Luna in the desktop app. Announced alongside prompt-caching improvements that let you change reasoning effort without invalidating the cache. $2/$10 (Sol) · $0.10/$0.50 (Luna)
In its own press release at Apsara Conference 2026 (Sep 22, Hangzhou), Alibaba said only that its next-generation model, Qwen 4, is currently in training. No release date, parameter count, licence, price or API identifier has been published. The same release sets out a roadmap in which the following Qwen 4.5 and Qwen 5 series are projected to scale to 5–10 trillion parameters (that is not a Qwen 4 specification). Also announced: Qwen3.8-LiveTranslate, the Qwen-Audio-3.1 family, the Qwen Intelligence agent platform for phones, and Qwen-Image 3.1 later in 2026.
How we treat it: announced but not shipped, so no dedicated page — recorded here, and we will publish pricing and specs once confirmed in primary sources.Price not published
The second model in Anthropic's Claude 5.5 family; the docs call it "the best combination of speed and intelligence". $2 input / $10 output / $0.20 cache reads / $2.50 5-minute cache writes / $4 one-hour cache writes — identical to the previous-generation Sonnet 5, not a price cut. 1M context, 128K max output, knowledge cutoff June 2026. Adaptive thinking is on by default with a default effort of high, and between_tools turns off up-front thinking. No benchmark table is published. Available on the Claude API, Amazon Bedrock, Vertex AI, Microsoft Foundry and Claude Platform on AWS. $2/$10 (cache read $0.20, write $2.50)
A new frontier model.
Release stages: (1) rolling out now through the Fairwind Program — the post's term is "trusted cyber defenders"; a limited program for invited testers such as security specialists, not an open sign-up; (2) to developers, enterprises and consumers "as soon as possible", starting with paid API customers and Google AI Ultra subscribers per the official post. Google says it is engaged in the U.S. government's voluntary pre-release access process while it iterates on guardrails and expands access.
What you can do today: the only route is the Fairwind Program (invitation-based) — no general API or consumer access yet, and you cannot buy at this price yet. Context: 1 million tokens (per the post summary).
Pricing: $2 per 1M input tokens / $10 per 1M output tokens, with cached input at 95% off the input price, is announced as an introductory price for launch; the post does not state that billing has started.
Benchmarks: the official post publishes no benchmark table, so we do not list aggregate-site scores as numbers here (any mention stays at the level of "as reported by ...").Announced intro price $2/$10 (cached input 95% off)