Major LLM Model Release Timeline [Jan 2025 – Sep 2026]

Release timeline of major OpenAI (GPT-5), Anthropic (Claude 4/5), Google (Gemini 3), DeepSeek (R1/V4) and xAI (Grok) models, verified against each company's official announcements. See Pricing Comparison for current rates and the Token Calculator for cost estimates.

●Last updated: | Primary-source links per entry
OpenAIAnthropicGoogleDeepSeekZhipu/Z.aixAI
2025
2025-01-20 DeepSeek

DeepSeek R1 released — the origin of the "DeepSeek Shock"

Open-weight reasoning model released under the MIT license. Achieved OpenAI o1-class reasoning at low cost, triggering the "DeepSeek Shock."

2025-08-07 OpenAI

GPT-5 released — router-based system unifying fast response and reasoning

OpenAI's first unified system that auto-selects between a fast-response model and a reasoning model via a real-time router. Current generation is the GPT-5.6 family.

2025-11-18 Google

Gemini 3 Pro released — 1M-token multimodal understanding tops LMArena

1M-token context window and enhanced multimodal understanding. Recorded 1501 Elo on LMArena.

2026
2026-02-05 Anthropic

Claude Opus 4.6 released — 1M-token context on the flagship

Introduced a 1M-token context window to Anthropic's flagship. Long-document retrieval accuracy significantly improved.

2026-02-17 Anthropic

Claude Sonnet 4.6 released

Balanced model with upgraded coding, computer use, long-form reasoning and agentic planning. $3.00/$15.00

2026-02-19 Google

Gemini 3.1 Pro Preview released

Major abstract reasoning gains. 77.1% on ARC-AGI-2. $2.00/$12.00 (≤200K)

2026-03-05 OpenAI

GPT-5.4 released — native Computer Use API debut

Native Computer Use (OSWorld 75%). Five-level reasoning effort control. $2.50/$15.00

2026-04-16 Anthropic

Claude Opus 4.7 released — new xhigh reasoning effort

Coding (SWE-bench Pro 64.3%) and vision gains. Began adopting a new tokenizer.

2026-04-23 OpenAI

GPT-5.5 released

Then-flagship with stronger agentic computer use and scientific research support. $5.00/$30.00

2026-04-24 DeepSeek

DeepSeek V4 Pro/Flash Preview released

1M-token context, MIT license. V4 Pro (1.6T/49B active) · V4 Flash (284B/13B active). V4 Pro $0.435/$0.87 (promo at the time, now retired)

2026-05-19 Google

Gemini 3.5 Flash GA — Google I/O 2026

Optimized for agents and coding. Beats 3.1 Pro benchmarks at 4× speed. $1.50/$9.00

2026-05-28 Anthropic

Claude Opus 4.8 released

Code-defect misses reduced to roughly 1/4. Dynamic Workflows research preview. $5.00/$25.00

2026-06-09 Anthropic

Claude Fable 5 / Mythos 5 released — new "Mythos class" tier

New tier above Opus. Fable 5 general availability, Mythos 5 limited. Paused 6/12 → restored 7/1. Fable 5 $10.00/$50.00

2026-06-30 Anthropic

Claude Sonnet 5 released

The "most agentic" Sonnet. Near-Opus 4.8 performance at a low price. The $2/$10 introductory price was made permanent on 2026-08-10 (the $3/$15 increase scheduled for Sep 1 will not occur). $2.00/$10.00 (permanent)

2026-07-09 OpenAI

GPT-5.6 family released — Sol / Terra / Luna

New flagship Sol, balanced Terra, and fastest low-cost Luna released together. Terra/Luna price cuts on 7/30. Sol $5/$30 · Terra $2/$12 · Luna $0.20/$1.20 (after 7/30)

2026-07-21 Google

Gemini 3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber released New

3.6 Flash has cheaper output than 3.5 Flash ($7.50 vs $9.00). 3.5 Flash-Lite is the cheapest Gemini 3.x ($0.30/$2.50) for high-throughput agents. 3.5 Flash Cyber is cybersecurity-focused. 3.6 Flash $1.50/$7.50 · Flash-Lite $0.30/$2.50

2026-07-24 Anthropic

Claude Opus 5 released — Opus 4.8 successor, same price, better performance New

Premium successor to Opus 4.8. Same price ($5/$25) with improved coding and professional tasks. Fast mode ($10/$50). ~85% fewer safety-classifier interventions vs Fable 5. $5.00/$25.00

2026-07-30 OpenAI

GPT-5.6 Terra/Luna major price cuts Updated

Terra: $2.50→$2.00 / $15.00→$12.00. Luna: $1.00→$0.20 / $6.00→$1.20 (80% cut). Priority processing → Fast mode renamed the same day. Terra $2/$12 · Luna $0.20/$1.20

2026-07-31 DeepSeek

DeepSeek V4 Flash officially released (public beta) New

API model name deepseek-v4-flash. Native Responses API and Codex support. Agent benchmarks far exceed V4-Pro-Preview. V4-Pro API and APP/WEB unchanged.

2026-08-12 xAI

Grok 4.6 released — new flagship for code & agents

Successor to Grok 4.5. Scores 61 on the AA Intelligence Index, matching GPT-5.6 Sol. 500K context, knowledge cutoff Feb 1 2026. $2/$6 (fast 2×)

2026-08-13 DeepSeek

DeepSeek V4 Pro GA released — peak/off-peak pricing introduced New

GA rollout on APP, Web and API (model name deepseek-v4-pro). Responses API support, three thinking-effort levels (low/high/max). Peak/off-peak pricing from 8/16 16:00 UTC (off-peak = half of peak). Weekends off-peak all day since 8/23. V4 Pro output off-peak $1.98 / peak $3.96

2026-08-14 Google

Gemini 3.7 Flash released — most capable Flash for coding & agents

Three weeks after 3.6 Flash. Intro price at half ($0.75/$3.75 thru 12/31 → $1.50/$7.50). FrontierCode 43.6% · DeepSWE 65.3%. $0.75/$3.75 (thru 12/31)

2026-08-19 OpenAI / Cloudflare

GPT-5.6 Sol 50% off via Cloudflare AI Gateway New

Routing to openai/gpt-5.6-sol through Cloudflare AI Gateway applies 50% off automatically (Unified Billing only; Bring Your Own Keys excluded). Input $2.50 / output $15 / cache read $0.25 (through 2026-09-18).

2026-08-21 DeepSeek

DeepSeek V4 Flash Vision Exp released (experimental) New

Experimental multimodal vision model, accessed via model='deepseek-v4-flash-vision-exp'. On par with V4 Flash on pure text; large gains on vision agent benchmarks.

2026-08-21 OpenAI

GPT-5.6 Sol promo price — $4/$20 (at least through 11/21) New

Flagship Sol cut from $5→$4 input / $30→$20 output for 3 months. Applies to Fast mode, long-context, and Batch/Flex too. Was $5/$30. Promo guaranteed "at least through 2026-11-21".$4/$20 (Short context)

2026-08-26 Zhipu / Z.ai

GLM 5.3 Flash released — natively multimodal at Flash cost New

The first natively multimodal GLM-5 model. 320B/18B MoE, 1M context, 128K output, MIT. Beats GLM-5.2 on coding/agentic benchmarks at ~1/10 the price. $0.15/$0.50 (list)

2026-08-26 Alibaba / Qwen

Qwen3.8-Flash released — low-cost multimodal, Qwen4 architecture preview New

A 125B (6B active) + 51B N-gram multimodal MoE. The GDN+QSA hybrid beats Qwen3.7-Plus on coding, office and agentic benchmarks at ~1/9 the training cost. Open-weight Flash-Next also on HF/ModelScope. $0.15/$0.47 (international)

2026-09-01 Anthropic

Claude Fable 5.1 / Mythos 5.1 released — new Mythos-class flagship New

A new frontier for coding, knowledge work and long-horizon agents. Same model, different safeguards (Fable 5.1 generally available / Mythos 5.1 trusted-access only). Cache reads cut 75% ($1.00→$0.25). 1M context, 128K output. Terminal-Bench-Science 52.6% (more than double Fable 5's 24.7%). $10/$50 · cache read $0.25

2026-09-02 Google

Gemini 3.8 Flash / 3.8 Flash Cyber released — best reasoning & coding Flash New

Three weeks after 3.7 Flash — third Flash release in six weeks. Intro $0.75/$3.75 (thru 12/31 → $1.50/$7.50). DeepSWE 73.8% · HLE-Verified 54.9%. Cybersecurity-focused 3.8 Flash Cyber is Fairwind Program-only. $0.75/$3.75 (thru 12/31)

2026-09-03 OpenAI

GPT-6 Astra released — new-generation flagship ($10/$50) New

Successor to GPT-5.6 Sol. State-of-the-art computer use, coding, cybersecurity and science. ARC-AGI-3 99.9% · ExploitBench 100% · Terminal-Bench Science 64.6%. 1.05M context, 128K output. Staged rollout (Daybreak → API/Plus). $10/$50

2026-09-10 DeepSeek

DeepSeek-V4.1-Flash released — steep price cuts, native image input (V4 Pro routing later withdrawn) New

The smallest model in the new architecture family: 1M context, 384K output, native image input, thinking by default. API name deepseek-flash. V4 Flash and V4 Flash Vision Exp are retired and consolidated into V4.1 Flash. DeepSeek initially said V4 Pro requests would route to V4.1 Flash from 2026-09-14, but that notice was later withdrawn and V4 Pro continues (no V4.1 Pro release date announced). $0.15/$0.60 (off-peak, cache-miss input/output)

2026-09-13 DeepSeek (verified by this site)

DeepSeek V4 Pro confirmed to continue — the 9/14 routing notice is withdrawn

Footnote (2) of the official pricing page and the change log now read: "In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." The V4 Pro price row (off-peak input $0.66 / output $1.98) is retained, so deepseek-v4-pro keeps its own rates. $0.66/$1.98 (off-peak, cache-miss input/output)

2026-09-21 xAI

Grok 4.7 released — larger base model, $2/$6 pricing unchanged New

Successor to Grok 4.6. 500K context, May 2026 knowledge cutoff. Beats Grok 4.6 on all seven official benchmarks and leads the four-model field on EEBench and the Harvey Legal Agent Benchmark. Third-party testing puts it at 46 on the Artificial Analysis Intelligence Index versus 53 for Fable 5.1 and GPT-6. $2/$6 (Fast 2x · $4/$12 above 200K)

2026-09-22 Anthropic

Claude Opus 5.5 released — first of the Claude 5.5 family, cheaper than Opus 5 New

Anthropic's first model in the new Claude 5.5 family. The official claim: performance at the level of Fable 5.1 on most work at 40% lower cost than Opus 5 on typical workloads. Input $5→$4, output $25→$20, cache reads $0.50→$0.20 (down 60%); Fast mode $8/$40. 1M context, 128K max output, June 2026 knowledge cutoff. Thinking can no longer be disabled — depth is set with effort (default medium). Sonnet 5.5 shipped on 9/28 (see the entry below); Haiku 5.5 is still unconfirmed. $4/$20 (cache read $0.20, write $5)

2026-09-22 Xiaomi

Xiaomi MiMo-V2.6 released — natively omnimodal, 1M context, pricing unchanged from V2.5 New

Three variants: Pro, Flash and Pro-UltraSpeed. Its official model page confirms text, image, video and audio input, text output and a 1M-token context. Pricing: Flash $0.14/$0.28, Pro $0.435/$0.87, Pro-UltraSpeed $4.35/$8.7 (cache hits $0.0028 / $0.0036 / $0.036). Xiaomi claims 46.32 on the Artificial Analysis Intelligence Index v4.3 and calls it “the strongest open-source model to date”. Technical report, training environment and RL code are released, with the production RL run live-streamed. $0.14 / $0.435 / $4.35 (input per 1M)

2026-09-22 OpenAI

GPT-6 Sol and GPT-6 Luna released — 50% cheaper than the GPT-5.6 versions New

The two practical tiers of the GPT-6 generation. Sol falls from $4/$20 to $2/$10; Luna from $0.20/$1.20 to $0.10/$0.50. Cached input is discounted 90% ($0.20 for Sol, $0.01 for Luna). 1,050,000-token context, 128K max output. GPT-5.6 Sol and Luna become legacy. Available in ChatGPT Work and Codex; Free and Go users get Luna in the desktop app. Announced alongside prompt-caching improvements that let you change reasoning effort without invalidating the cache. $2/$10 (Sol) · $0.10/$0.50 (Luna)

2026-09-22 Alibaba / Qwen

Qwen 4 announced as “in training” — no date, price or weights New

In its own press release at Apsara Conference 2026 (Sep 22, Hangzhou), Alibaba said only that its next-generation model, Qwen 4, is currently in training. No release date, parameter count, licence, price or API identifier has been published. The same release sets out a roadmap in which the following Qwen 4.5 and Qwen 5 series are projected to scale to 5–10 trillion parameters (that is not a Qwen 4 specification). Also announced: Qwen3.8-LiveTranslate, the Qwen-Audio-3.1 family, the Qwen Intelligence agent platform for phones, and Qwen-Image 3.1 later in 2026.
How we treat it: announced but not shipped, so no dedicated page — recorded here, and we will publish pricing and specs once confirmed in primary sources.Price not published

2026-09-28 Anthropic

Claude Sonnet 5.5 released — the balanced model of the Claude 5.5 family, at the same price as Sonnet 5 New

The second model in Anthropic's Claude 5.5 family; the docs call it "the best combination of speed and intelligence". $2 input / $10 output / $0.20 cache reads / $2.50 5-minute cache writes / $4 one-hour cache writes — identical to the previous-generation Sonnet 5, not a price cut. 1M context, 128K max output, knowledge cutoff June 2026. Adaptive thinking is on by default with a default effort of high, and between_tools turns off up-front thinking. No benchmark table is published. Available on the Claude API, Amazon Bedrock, Vertex AI, Microsoft Foundry and Claude Platform on AWS. $2/$10 (cache read $0.20, write $2.50)

2026-09-30 Google

Gemini 4 Argon announced — rolling out to trusted cyber defenders first, general access phased New

A new frontier model.
Release stages: (1) rolling out now through the Fairwind Program — the post's term is "trusted cyber defenders"; a limited program for invited testers such as security specialists, not an open sign-up; (2) to developers, enterprises and consumers "as soon as possible", starting with paid API customers and Google AI Ultra subscribers per the official post. Google says it is engaged in the U.S. government's voluntary pre-release access process while it iterates on guardrails and expands access.
What you can do today: the only route is the Fairwind Program (invitation-based) — no general API or consumer access yet, and you cannot buy at this price yet. Context: 1 million tokens (per the post summary).
Pricing: $2 per 1M input tokens / $10 per 1M output tokens, with cached input at 95% off the input price, is announced as an introductory price for launch; the post does not state that billing has started.
Benchmarks: the official post publishes no benchmark table, so we do not list aggregate-site scores as numbers here (any mention stays at the level of "as reported by ...").Announced intro price $2/$10 (cached input 95% off)

About this timeline

  • Inclusion criteria: Only models verified against official company blogs / release notes as primary sources.
  • Scope: Limited to flagship-tier and high-profile releases from the five major providers.
  • Pricing: See the Pricing Comparison for current rates (prices shown here are simplified).
  • Update policy: Appended as new models are released.

⚠️ Disclaimer

  • Release dates are confirmed via official announcements or multiple independent sources. Omissions or errors are possible.
  • Pricing changes frequently. Always check the Pricing Comparison and official pages before contracting.