The short version
- It replaces Opus 5; it is not a cut-down Fable 5.1. Anthropic says it matches Fable 5.1 on most work and is 40% cheaper than Opus 5 on typical workloads.
- The price cut is the headline. Input $5→$4, output $25→$20, cache reads $0.50→$0.20 (a 60% cut). Anthropic says cache reads make up the majority of agentic and coding work costs.
- Thinking can no longer be disabled. Depth is controlled with an
effortparameter (low / medium / high…). The default is medium — the announcement's chart notes say “at default effort (medium)”. - Sonnet 5.5 was released on 2026-09-28 — see our page for it. Haiku 5.5 is still described only as following "in the coming weeks"; as of 2026-10-01 we have not verified its availability.
Pricing (USD per 1M tokens)
Taken from the pricing table in Anthropic's announcement and the official pricing documentation — the same standard, pay-per-token basis we use on our pricing comparison.
| Item | Claude Opus 5.5 | Claude Opus 5 | Fable 5.1 (for reference) |
|---|---|---|---|
| Input | $4.00 | $5.00 | $10.00 |
| Output | $20.00 | $25.00 | $50.00 |
| Cache read (cache hit) | $0.20 | $0.50 | $0.25 |
| Cache write (5-minute TTL) | $5.00 | $6.25 | — |
Note the nuance: on Fable 5.1 a cache hit costs 2.5% of the standard input price ($0.25); on Opus 5.5 it costs 5% ($0.20). Opus 5.5 has the cheaper headline input rate, while Fable 5.1 keeps the deeper cache discount.
Fast mode (Claude Code and the Claude Platform)
| Mode | Input | Output | Speed |
|---|---|---|---|
| Standard | $4.00 | $20.00 | baseline |
| Fast mode | $8.00 | $40.00 | up to 2.5x |
Fast mode pricing “applies across the full context window, including requests over 200k input tokens” per Anthropic's pricing docs.
Primary sources: anthropic.com/claude-opus-5-5 (Pricing table) · platform.claude.com pricing docs (verified 2026-09-25)
Figure: how the rates line up
← scroll horizontally →
Specifications
Sources: Claude Platform Docs, Models overview (context window 1M, max output 128K, reliable knowledge cutoff Jun 2026, training data cutoff Jun 2026) · Anthropic announcement (availability, thinking behaviour).
Nine official benchmarks — and who actually ran each harness
From the table in Anthropic's announcement. Unless otherwise noted, every Opus 5.5 figure uses adaptive thinking at max effort (verbatim: “Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort”).
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 (Main) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 | 57.8% | 51.8% | 46.6% | — | 41.7% |
| GDPval-AA v2.1 (Elo) | 1,846 | 1,735 | 1,708 | 1,542 | 1,588 |
| AutomationBench (business workflows) | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity's Last Exam (with tools) | 67.7% | 65.6% | 63.6% | 57.2% | — |
| Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| OSWorld 2.0 (computer use, partial) | 81.8% | 80.7% | 74.0% | — | — |
| Chartography (with tools) | 89.0% | 88.4% | 83.4% | — | — |
Who measured what (this is the important part)
| Benchmark | Run and reported by | Harness and conditions |
|---|---|---|
| Terminal-Bench 4.0 | Anthropic's own setup, plus the public leaderboard | Public leaderboard uses the Claude Code harness, 5 trials/task, and reports Opus 5 at 51.8%; Anthropic's setup reproduces it at 52.3% (within noise). Opus 5.5 was run at xhigh effort and GPT-6 Astra at high effort — each model's highest score. Standard error ±2.6 pts. GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI. |
| Terminal-Bench-Science 0.1 | Anthropic's setup, plus the public leaderboard | Public leaderboard uses 3 trials/task and the Claude Code harness, reporting Opus 5 at 30.0% versus 29.0% on Anthropic's setup. Standard error ±3.5–5 pts. GPT-6 Astra figure is as reported by OpenAI. |
| AutomationBench | Zapier ran and reported it | Runs were performed without fallback models, so safeguard interventions counted as failures — a condition that understates the real-world score. Opus 5.5 comes from Zapier's early-access evaluation; Opus 5, GPT-5.6 Sol and GPT-6 Astra come from Zapier's public leaderboard. |
| GDPval-AA v2.1 | Artificial Analysis' evaluation, quoted by Anthropic | Anthropic describes it as “Artificial Analysis's GDPval-AA v2.1 … real-world professional work across 44 occupations”, scored in Elo. |
| CursorBench 4.0 / FrontierCode v1.1 | Anthropic's own setup | Chart notes also give figures “at default effort (medium)” — FrontierCode 54.6%, CursorBench 52.5%. So the table shows max-effort scores while the charts show default-effort scores: different conditions. |
| WANDR (chart only, not in the table) | Perplexity's benchmark, run by Anthropic under modified conditions | Verbatim note: “Claude models were run with offline versions of the web search and web fetch tools, programmatic tool calling, code execution, and a 980k-token task budget. This differs from Perplexity's published setup, scores are not directly comparable”. Only models scored under identical conditions are shown. |
Source: Anthropic, “Introducing Claude Opus 5.5” (benchmark table and footnotes 1–4, verified 2026-09-25)
What the efficiency claims mean
| Comparison | What the announcement says |
|---|---|
| Opus 5.5 vs Opus 5 | 40% lower cost on typical workloads; output generated more than 30% faster. |
| Terminal-Bench 4.0 | Opus 5.5 at default effort beats Opus 5 at max effort for about one fifth of the cost, and matches GPT-6 Astra at about 40% of the cost. |
| CursorBench 4.0 | Beats GPT-5.6 Sol's top score by 11 points for about a third of the cost. |
| Internal test: 200,000-line codebase audit | Opus 5.5 finished in under three hours; Opus 5 took over 20 hours and used 2.5x as many tokens. |
| Internal test: HAProxy C→Rust rewrite | Opus 5.5 finished in 9.5 hours versus 12 for Fable 5.1, and cost 51% less. |
All of these are vendor-run internal tests, not independent reproductions.
Safety: safeguards, and a stated limit of the evidence
| Item | What the announcement says |
|---|---|
| External evaluation | Tested before release by Frontier Design and METR. |
| Cybersecurity | Same class of safeguards as Fable 5.1; most cybersecurity tasks are re-routed to Opus 4.8, transparently. |
| Biology | Same biology safeguards as Fable 5.1. Vetted organisations can apply to the Life Sciences Verification Program. |
| Distillation | Launches with preserved thinking, the anti-distillation safeguard introduced with Fable 5.1. Applies to API accounts created on or after 2026-08-31. |
| New containment test | Attempted to circumvent boundaries about 85% less often than Opus 5 or Mythos 5.1; every attempt was low severity and self-reported. |
| Data handling | Available with zero data retention; ships with watermarking to comply with the EU AI Act. |
Source: Anthropic, “Introducing Claude Opus 5.5” (Safety section)
How this relates to our other Claude pages
| Model | On this site | Relationship |
|---|---|---|
| Claude Fable 5.1 | Dedicated page | A different model (higher tier, higher price) at $10 input / $50 output. Opus 5.5's pitch is Fable-5.1-level work at a lower cost. |
| Claude Mythos 5.1 | Dedicated page | A different model (trusted-access tier for vetted organisations). Opus 5.5's cyber and biology safeguards are the same class, and lifting them needs a verification program. |
| Claude Opus 5.5 | This page (new) | First model of the Claude 5.5 family; successor to Opus 5. Sonnet 5.5 was released 2026-09-28 (dedicated page). Haiku 5.5 availability is not verified. |
We have no page for /en/claude-opus-5/ as of 2026-10-01. Sonnet 5.5 is covered at /en/claude-sonnet-5-5/; we do not link to pages that do not exist.
Where it fits — and where to be careful
- Long-running agentic coding: cache reads at $0.20 and Anthropic's own note that cache reads dominate agentic costs make this a good fit for setups that resend repository context.
- Plan for thinking you cannot switch off: easy tasks still consume output tokens. Lower
effortor use a lighter model for simple work. - Cyber and biology work: requests transparently fall back to Opus 4.8 or Opus 5, so a task may be answered by a different model than the benchmark numbers describe.
- If raw peak accuracy is the goal: GPT-6 Astra leads Terminal-Bench-Science at 64.6% versus 58.7%. Check per domain rather than by overall reputation.
- The start of a price war: OpenAI announced GPT-6 Sol and Luna roughly a week later at about half the previous prices — see our GPT-6 Sol / Luna page.
Estimate your real cost
Compare token costs across models, including Claude Opus 5.5.
🧮 Open the token calculator →📊 Full pricing comparison · 🆚 GPT-6 Sol / Luna · 🤖 Fable 5.1 · 🗓️ Release timeline
⚠️ Disclaimer
- Pricing and specifications were verified on Anthropic's announcement and official documentation on 2026-09-25 and may change without notice. Always confirm on the official pages before you commit.
- Benchmark figures include vendor-reported numbers. GPT-6 Astra and GPT-5.6 Sol figures come from OpenAI's own reporting, not from Anthropic re-runs, and conditions may not be identical.
- This site is for information only and is not a recommendation or agent for any provider.