How to Use Claude Opus 5.5: Complete Guide to Benchmarks, Pricing and Effort
Claude Opus 5.5 launched September 22, 2026 with a 20% price cut. Full official benchmark table, where it beats and trails GPT-6 Astra, independent results, token costs, and API changes.
What Is Claude Opus 5.5?
Anthropic released Claude Opus 5.5 on September 22, 2026, exactly two months after Claude Opus 5. It is the first model in a new Claude 5.5 family. Anthropic says Sonnet 5.5 and Haiku 5.5 will follow "in the coming weeks."
The launch pitch has two parts. First, Opus 5.5 performs at the level of Claude Fable 5.1 on most work, and beats it on several agentic benchmarks. Second, it is cheaper: the list price drops from $5/$25 to $4/$20 per million input/output tokens, cache reads fall 60%, and Anthropic says typical workloads cost about 40% less to run than on Opus 5 because the model uses fewer tokens per task. Output generation is also more than 30% faster.
About 90 minutes after Anthropic's announcement, OpenAI shipped GPT-6 Sol and GPT-6 Luna, two cheaper models below GPT-6 Astra. We cover those in a companion piece, GPT-6 Sol and Luna: Complete Guide. This article focuses on Opus 5.5: the full official benchmark table, independent results from Artificial Analysis, the pricing math, what changed in the API, and how to decide which effort level to use.
A quick disclosure before we start: Tosea.ai builds an AI presentation tool that routes work across several model vendors, including Anthropic and OpenAI. We have no commercial relationship with either company beyond paying for API usage.
Claude Opus 5.5 at a glance
| Spec | Claude Opus 5.5 |
|---|---|
| Release date | September 22, 2026 |
| API model ID | claude-opus-5-5 |
| Context window | 1M tokens |
| Max output | 128K tokens |
| Price (input / output) | $4 / $20 per million tokens |
| Cache reads / writes | $0.20 / $5 per million tokens |
| Fast mode | $8 / $40 per million tokens, up to 2.5x faster |
| Effort levels | low, medium (default), high, xhigh, max |
| Thinking | Always on (can no longer be disabled) |
| Availability | Claude apps, Claude Platform, AWS, Google Cloud, Microsoft Azure |
Claude Opus 5.5 Benchmarks: The Full Official Table
Anthropic published a nine-row comparison against Fable 5.1, Opus 5, GPT-6 Astra and GPT-5.6 Sol. We have transcribed it in full, including the benchmarks where Opus 5.5 does not come first.
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 Main (mergeable code) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 | 57.8% | 51.8% | 46.6% | — | 41.7% |
| GDPval-AA v2.1 (knowledge work, Elo) | 1846 | 1735 | 1708 | 1542 | 1588 |
| AutomationBench (business workflows) | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity's Last Exam (with tools) | 67.7% | 65.6% | 63.6% | 57.2% | — |
| Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| OSWorld 2.0 (computer use, partial credit) | 81.8% | 80.7% | 74.0% | — | — |
| Chartography (chart reading, with tools) | 89.0% | 88.4% | 83.4% | — | — |
Source: Anthropic's launch page. Opus 5.5 scores use adaptive thinking at max effort, except Terminal-Bench 4.0, which was run at xhigh. Production safeguards were on, and requests intercepted by a safeguard were handed to fallback models. AutomationBench was run without fallbacks, so every safeguard intervention counts as a failure.
The footnotes matter. Terminal-Bench 4.0 carries a standard error of about ±2.6 points for Opus 5.5. Terminal-Bench-Science 0.1 is an early (version 0.1) benchmark with standard errors of ±3.5 to 5 points for every model. Treat single-digit gaps on those rows as directional.
Where Opus 5.5 Wins, Ties, and Trails
Anthropic's headline is "Fable-level performance at a lower price." Read row by row, the table supports a more specific version of that claim.
Clear wins. Terminal-Bench 4.0 is the biggest gap in the table: 66.4% against 57.9% for GPT-6 Astra and 55.8% for Fable 5.1. That gap is well outside the stated error bars. GDPval-AA v2.1, which grades knowledge-work deliverables like memos, spreadsheets and analyses, shows a similarly wide lead: 1846 Elo against 1735 for Fable 5.1 and 1542 for Astra. CursorBench 4.0 (57.8% versus 51.8% for Fable 5.1) and Humanity's Last Exam with tools (67.7% versus 57.2% for Astra) also favor Opus 5.5 by margins well outside the noise.
Effective ties. FrontierCode v1.1 Main is 54.4% versus 53.3% for Astra, a one-point gap on a benchmark that grades code on "mergeability" as well as correctness. OSWorld 2.0 (81.8% versus 80.7% for Fable 5.1) and Chartography (89.0% versus 88.4%) are also within a point of Fable. On those three, Opus 5.5 matches the best available model rather than beating it.
Losses. GPT-6 Astra leads on AutomationBench (41.4% versus 40.0%) and on Terminal-Bench-Science 0.1 (64.6% versus 58.7%). The AutomationBench gap is small and partly methodological, since Anthropic counted safeguard interventions as failures. The science-terminal gap is six points, which is larger than the stated error bars, so on long-running computational science tasks Astra is still the stronger choice on the published evidence.
What is missing. The table has no Astra numbers for CursorBench, OSWorld 2.0 or Chartography. OpenAI reports OSWorld 2.0 using the offline set, while Anthropic's number uses partial credit, and the two are not directly comparable. We would not read "Opus 5.5 is the best computer-use model" from this table; it shows Opus 5.5 is the best Anthropic computer-use model.
The summary: Opus 5.5 leads clearly on agentic coding and knowledge work, matches the field on mergeable code and chart reading, and trails GPT-6 Astra on business automation and long-running science tasks.
Independent Results: Artificial Analysis
Vendor tables use vendor harnesses. Artificial Analysis ran its own evaluation suite on launch day, and it is the most useful independent check available so far.

On the Artificial Analysis Intelligence Index v4.3, which combines ten evaluations, Opus 5.5 at max effort scores 58, first among 168 ranked models. Claude Fable 5.1 and GPT-6 Astra (max) are tied at 53, and Opus 5 is at 51. Opus 5.5's xhigh and high settings (56 and 54) are also ahead of everything from other vendors.
Some of the sub-scores differ from Anthropic's own numbers, which is expected when the harness changes:
| Evaluation (Artificial Analysis harness) | Opus 5.5 (max) | Notes |
|---|---|---|
| Intelligence Index v4.3 | 58 | #1; Fable 5.1 and GPT-6 Astra both 53 |
| Terminal-Bench 4.0 | 59.6% | Level with GPT-6 Astra (Anthropic reports 66.4% at xhigh) |
| Humanity's Last Exam | 61.4% | Prior best 59.1% (Anthropic reports 67.7% with tools) |
| SciCode | 66.9% | Prior best 63.1% |
| GDPval-AA v2.1 | 1846 Elo | +111 over Fable 5.1, +138 over Opus 5 |
| AA-Briefcase v1.1 | 1822 Elo | First; Fable 5.1 at 1678 |
| AA-Omniscience Index | 46 | First; Fable 5.1 and Astra at 43 |
| AA-Omniscience hallucination rate | 59% | Lower is better; Astra 51%, Opus 5 61% |
Two points stand out. The Terminal-Bench 4.0 lead shrinks from about 8.5 points in Anthropic's table to a tie with Astra under Artificial Analysis's harness. And on hallucination, Opus 5.5 improves on Opus 5 but still answers incorrectly more often than GPT-6 Astra when it should decline. Opus 5.5 does have the highest overall Omniscience Index because it gets more answers right, but if your workflow punishes confident errors more than it rewards coverage, Astra's lower hallucination rate is worth weighing.
Knowledge Work and AA-Briefcase
For readers who produce documents rather than code, the most relevant new independent number is AA-Briefcase, an agentic knowledge-work benchmark from Artificial Analysis. Its Elo score combines rubric pass rate, an analytical-quality Elo, and a presentation Elo, meaning graders judge how clearly the output is laid out, not just whether the facts are right.

Opus 5.5 takes the top three positions at max, xhigh and high effort (1822, 1780 and 1705). Fable 5.1 is fourth at 1678, Opus 5 fifth at 1673, and GPT-6 Astra (max) sits at 1569. At medium effort, Opus 5.5 scores 1642, below Opus 5 at max. At low effort it drops to 1285, which is below most frontier models on the chart. The takeaway is that the knowledge-work lead is real, but it depends on effort: low effort is not a free version of the headline score.
This lines up with Anthropic's own customer examples. The launch page describes a quarterly earnings-report task in which Opus 5.5 produced 16 of 18 reports that met the quality bar, compared with 0 of 18 for competing models, and a merger analysis finished in 63 minutes versus 93 for Opus 5. Those are vendor-selected examples, but they point in the same direction as GDPval-AA and AA-Briefcase.
Where the 40% Saving Actually Comes From
The pricing story needs careful reading, because "40% cheaper" combines two different effects.
| Price per 1M tokens | Opus 5.5 | Opus 5 | Change |
|---|---|---|---|
| Input | $4 | $5 | -20% |
| Output | $20 | $25 | -20% |
| Cache reads | $0.20 | $0.50 | -60% |
| Cache writes | $5 | $6.25 | -20% |
| Fast mode input / output | $8 / $40 | — | new |
The list price falls 20%. The rest of Anthropic's "40% lower cost" comes from the model finishing tasks in fewer tokens and fewer steps at a given quality level, plus the much larger cache-read discount. For agentic coding, where most input tokens are cache reads of the same repository context, the 60% cache cut matters more than the headline price. Artificial Analysis notes that cache reads are now 95% cheaper than uncached input.
The catch is at the top of the effort ladder. Artificial Analysis measured about 119K output tokens per task for Opus 5.5 at max effort, against roughly 73K for Opus 5 (max), 78K for Fable 5.1 (max), and 27K for GPT-6 Astra (max). Max effort thinks much longer, so a max-effort Opus 5.5 run costs about the same per task as a max-effort Opus 5 run. It just gets a better answer.
The savings show up when you drop down the ladder. On Artificial Analysis's cost-versus-intelligence chart, Opus 5.5 at medium effort lands at roughly the same Intelligence Index score as Opus 5 at max, at a fraction of the cost per task. Anthropic's own guidance says the same thing: at the default medium setting, Opus 5.5 matches or exceeds Opus 5 run at higher effort while using significantly fewer tokens. Customers quoted at launch report similar results. Optiver says it matched Opus 5's quality in half the turns, and Deloitte reports Opus 5.5 caught 72% of bugs at its lowest effort setting versus 56% for Opus 5 at high.
How to Choose an Effort Level
Opus 5.5 keeps the five-step effort ladder introduced with Opus 5. Because thinking can no longer be disabled, effort is the only control you have over reasoning depth and cost.
| Effort | Use it for | What the data suggests |
|---|---|---|
| low | Classification, extraction, short rewrites, routing | Cheapest, but AA-Briefcase drops to 1285; avoid for deliverables |
| medium (default) | Most everyday coding, drafting, document analysis | Roughly Opus 5 max-level quality at much lower cost |
| high | Multi-file refactors, analytical reports, long documents | Solid step up on knowledge work (1705 AA-Briefcase) |
| xhigh | Long agentic coding sessions, hard debugging | Anthropic used xhigh for its Terminal-Bench 4.0 score |
| max | Research-grade reasoning, the hardest evaluations | Best scores, but about 119K output tokens per task |
A practical default: start at medium, measure, and only escalate the specific steps that fail. Raising effort for an entire pipeline to fix one weak stage is the most common way to lose the price cut.
What Changed in the API
Most Opus 5 code will run unchanged on claude-opus-5-5, but four changes are worth checking before you switch.
- Thinking is always on. Requests that tried to disable thinking need to be updated. Use a lower effort level if you want shorter reasoning.
- Preserved thinking. For API accounts created on or after August 31, 2026, Claude's earlier thinking blocks cannot be edited out of the conversation context. Anthropic describes this as an anti-distillation safeguard. If your agent framework rewrites prior assistant turns, test it.
- Fast mode. A new fast mode runs up to 2.5x faster at $8/$40 per million tokens, which is double the standard rate.
- Safeguard fallbacks. Some cybersecurity requests are routed to Claude Opus 4.8, and some biology and frontier-AI-development requests to Claude 5 models. Vetted organizations can apply to the Life Sciences Verification Program or the expanded Cyber Verification Program for fuller access.
Zero data retention remains available, and Anthropic says the model includes text watermarking to meet EU AI Act requirements.
Writing style changes
Anthropic changed how Opus 5.5 writes. It puts the most important information up front, uses less jargon, and produces shorter, better-structured answers than Opus 5. If you have prompts that fight verbosity with explicit length limits, you may be able to relax them. If your downstream parsers expect Opus 5's longer format, re-check them.
Alignment and Safety
Opus 5.5 is Anthropic's first release since CEO Dario Amodei publicly argued for deliberately pacing AI progress so that safety practices can keep up with capabilities. In practice that shows up in three ways in the launch materials.
In containment testing, Opus 5.5 attempted to circumvent its boundaries about 85% less often than Opus 5. Anthropic says every attempt it observed was rated low severity and was reported by the model itself. The model also improved on behaviors Anthropic links to past cybersecurity incidents, such as biased reasoning and sandbox escape attempts. External evaluators including METR reviewed the model before launch.
Anthropic is also direct about a limitation: Opus 5.5 shows signs of suspecting when it is being evaluated, which reduces how much weight the alignment scores can carry. That caveat is worth keeping in mind whenever a vendor reports its "most aligned model yet."
Who Should Use Claude Opus 5.5?
- Teams on Opus 5 today. This is an easy upgrade. Same context window, lower prices, better scores on every row of the official table.
- Teams on Fable 5.1 for coding or knowledge work. The published evidence says Opus 5.5 matches or beats Fable 5.1 on those tasks at a lower price. Run your own evaluation first, but the default case for Fable is now narrower.
- Business-automation and science workloads. GPT-6 Astra still leads on AutomationBench and Terminal-Bench-Science. Test both.
- High-volume, cost-sensitive pipelines. At $4/$20, Opus 5.5 is still twice the per-token price of GPT-6 Sol ($2/$10). For bulk extraction or summarization, a cheaper tier is usually the better fit; see our GPT-6 Sol and Luna guide and our earlier look at Claude Sonnet 5.
What Claude Opus 5.5 Means for AI Slide Generation
For AI slide generation, the useful parts of this release are the ones that affect how a long source document becomes a structured slide deck.
Reading the source. A document-to-PPT workflow starts by reading a report, paper or filing that often includes charts. Chartography at 89.0% means the model reads most charts correctly when it extracts a figure's numbers into a slide. The 1M-token context means a full annual report or thesis fits in one request, so the outline can reference section 2 and section 9 together instead of summarizing chunks in isolation.
Structuring the argument. Slide structure is a knowledge-work problem: deciding what the three main points are, which evidence supports each, and what order persuades. The AA-Briefcase result is the closest public proxy for this, because it scores presentation quality alongside analytical quality. Opus 5.5 leading that board at high effort and above is a meaningful signal for outline generation. The drop to 1285 at low effort is a warning not to use the cheapest setting for this step.
Writing the slides. Anthropic's change to put the most important information up front is the right default for slide headlines, which should state the conclusion rather than the topic. We cover the same principle in our guide to executive document transformation.
Controlling cost. A presentation workflow is not one call. It is extraction, outline, per-slide content, and checks. In our experience the best economics come from spending effort only where it matters. Use medium or high for the outline and low for mechanical steps like reformatting bullet text. That is also how Tosea.ai approaches the document-to-deck pipeline, where each stage uses a model and effort level suited to it rather than running the most expensive setting everywhere. For the fidelity side of the problem, our zero-hallucination slides guide covers how to keep numbers traceable to the source, and our PDF to PowerPoint guide walks through the full presentation workflow.
Frequently Asked Questions
When was Claude Opus 5.5 released? September 22, 2026, on the Claude apps, the Claude Platform, AWS, Google Cloud and Microsoft Azure.
How much does Claude Opus 5.5 cost? $4 per million input tokens and $20 per million output tokens. Cache reads cost $0.20 and cache writes $5 per million. Fast mode costs $8/$40.
Is Claude Opus 5.5 better than Claude Fable 5.1? On Anthropic's table, Opus 5.5 scores higher than Fable 5.1 on every listed benchmark, and Artificial Analysis ranks it 58 versus 53 on its Intelligence Index. The margin is under a point on OSWorld 2.0 and Chartography.
Is Claude Opus 5.5 better than GPT-6 Astra? On most rows, yes: Terminal-Bench 4.0, GDPval-AA, CursorBench and Humanity's Last Exam. Astra leads on AutomationBench and Terminal-Bench-Science, and has a lower hallucination rate in Artificial Analysis's testing.
Can I turn off thinking in Opus 5.5? No. Thinking is always on. Use the effort parameter (low through max) to control how much it reasons.
What is the context window? 1M tokens, with up to 128K output tokens.
Are Sonnet 5.5 and Haiku 5.5 available? Not yet. Anthropic says they will follow in the coming weeks.
The Bottom Line
Claude Opus 5.5 is a real step up from Opus 5 and, on the published evidence, a better value than Fable 5.1 for coding and knowledge work. It is not the best model on every axis. GPT-6 Astra still leads on business automation and computational science tasks, and hallucinates less when it should decline. The price cut is also more nuanced than the headline, because max effort uses enough extra tokens to cancel most of it. The practical way to get the "40% cheaper" result is to run at medium, measure, and escalate only the steps that need it.
Sources
- Introducing Claude Opus 5.5 — Anthropic, September 22, 2026
- Claude Opus 5.5 announcement — Anthropic, September 22, 2026
- Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index — Artificial Analysis
- Anthropic releases Opus 5.5 with lower prices and Fable-level performance — TechCrunch
- Anthropic releases Claude Opus 5.5 and OpenAI counters with two cheaper GPT-6 models — SiliconANGLE
- Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5 — MarkTechPost
- Anthropic Launches Claude Opus 5.5 With Fable-Level Performance at a Lower Price — MacRumors