GuidesTosea Team16 MIN READ

How to Use Claude Opus 5: Complete Guide to Benchmarks, Effort and Pricing

Claude Opus 5 launched July 24, 2026 at Opus 4.8 pricing. Full official benchmark table, where it beats and loses to Fable 5 and GPT-5.6 Sol, the effort ladder, API changes, and migration notes.

How to Use Claude Opus 5: Complete Guide to Benchmarks, Effort and Pricing

Anthropic shipped Claude Opus 5 on July 24, 2026. The pitch is unusually specific: near-Fable-5 intelligence at half the price, on the same $5 / $25 per million tokens that Opus 4.8 already charged. No price increase, a step change in capability, and — for the first time in the Opus line — thinking turned on by default.

The interesting part is not the headline. It is that Anthropic published its results as cost-versus-score curves across five effort levels rather than a single number per benchmark, which turns "how good is this model" into "how good is it at the budget you are willing to spend." That reframing is the actual product, and it changes how you should wire Opus 5 into anything that runs at scale.

This guide covers the full official benchmark table, the benchmarks where Opus 5 loses, the effort ladder and what it costs, the two breaking API changes, and what all of it means if your workload is turning documents into presentations. For the previous generation, see our complete guide to Claude Opus 4.8; for the tier above it, our Claude Fable 5 guide.

What Is Claude Opus 5?

Claude Opus 5 (claude-opus-5 on the Claude API) is Anthropic's model for complex agentic coding and enterprise work. It sits below Claude Fable 5 in raw capability and above Claude Sonnet 5 in cost, and Anthropic positions it explicitly as the everyday workhorse rather than the trophy model: it is the new default on Claude Max and the strongest model available on Claude Pro.

The specifications:

PropertyClaude Opus 5
API model IDclaude-opus-5
Context window1M tokens (default and maximum)
Max output128K tokens
Pricing$5 / $25 per million input / output tokens
Fast mode~2.5× speed, $10 / $50, Claude API only
ThinkingOn by default (adaptive)
Effort levelslow, medium, high, xhigh, max
Prompt cache minimum512 tokens (down from 1,024)
Data retention requirementNone for general access

It is available on the Claude API, Amazon Bedrock (as anthropic.claude-opus-5), Google Cloud, and Microsoft Foundry. Unlike Fable 5, it carries no 30-day data-retention requirement for general access — which matters more than it sounds for regulated buyers who could not touch Fable 5 at all.

Claude Opus 5 Benchmarks: The Full Official Table

Anthropic published a head-to-head table against Fable 5, Opus 4.8, and OpenAI's GPT-5.6 Sol. This is the complete set, transcribed from the official results, including the rows where Opus 5 does not win:

Claude Opus 5 official benchmark table comparing Frontier-Bench, GDPval-AA v2, ARC-AGI-3, BrowseComp, Humanity's Last Exam, OSWorld 2.0, DeepSWE, FrontierCode, AutomationBench, legal, health and biology results against Fable 5, Opus 4.8 and GPT-5.6 Sol

BenchmarkOpus 5Fable 5Opus 4.8GPT-5.6 Sol
Frontier-Bench v0.1 (agentic terminal coding)43.3%33.7%21.1%34.4%
GDPval-AA v2 (knowledge work, Elo)1861174715931736
ARC-AGI-3 (novel problem-solving)30.2%1.5%7.8%
BrowseComp (agentic search)90.8%87.4%84.3%90.4%
Humanity's Last Exam (no tools)56.3%56.5%49.8%
Humanity's Last Exam (with tools)64.7%63.9%57.9%
OSWorld 2.0 (computer use)70.6%66.1%55.7%62.6%
DeepSWE v1.1 (agentic coding)68.8%69.7%59.0%72.7%
FrontierCode v1.1, Main53.4%53.5%46.5%47.5%
AutomationBench (business workflows)26.0%17.4%17.0%18.1%
Legal Agent Benchmark (held-out)11.7%13.3%10.4%2.5%
HealthBench Professional59.8%66.0% (Mythos 5)57.4%60.5%
BioMysteryBench (hard)49.4%46.5%42.4%
BioMysteryBench (human-solved)90.1%89.0% (Mythos 5)88.5%

Two methodology notes before anyone quotes these. The Frontier-Bench figure comes from an internal run on the mini-SWE-agent harness with a GKE backend, averaged over five attempts per task — and Opus 4.8 served as the fallback when safety classifiers refused a request for either Opus 5 or Fable 5. And the Fable 5 column silently becomes Mythos 5 on the health and biology human-solved rows, which are Anthropic's restricted-access model rather than the generally available one.

On the independent side, Artificial Analysis has Claude Opus 5 (adaptive reasoning, max effort) at 61 on its Intelligence Index — first out of 187 models, ahead of Fable 5 at 60 and GPT-5.6 Sol at 59. That index aggregates nine evaluations including GDPval-AA v2, Terminal-Bench v2.1, SciCode, GPQA Diamond and Humanity's Last Exam, so it is a broader read than any single benchmark. It comes with a caveat worth repeating: Artificial Analysis flags Opus 5 as unusually verbose, generating roughly 100M output tokens across its evaluation suite against a 63M average. On a per-token bill, verbosity is not free.

LMArena has not yet published a rating. Opus 5 landed a day ago and new models typically need one to two weeks of votes before their Elo stabilizes, so anyone citing an arena rank for Opus 5 this week is citing noise.

Where Opus 5 Wins — and Where It Loses

Eight of the fourteen rows above go to Opus 5, and the margins on those wins are not subtle. Frontier-Bench is the standout: 43.3% against Opus 4.8's 21.1% is more than double the score at a lower cost per task, and it clears GPT-5.6 Sol's 34.4% by nine points. ARC-AGI-3 is stranger still — 30.2% against 7.8% for GPT-5.6 Sol and 1.5% for Opus 4.8, roughly four times the next-best model on a benchmark specifically built around problems the model has never seen. AutomationBench shows a similar shape at 26.0% versus a field clustered between 17% and 18%.

The losses are worth stating just as plainly, because they cluster in a revealing way:

  • DeepSWE v1.1 — 68.8% vs GPT-5.6 Sol's 72.7%. This is the clearest defeat in the table, and it is on agentic coding, the exact territory Opus 5 is marketed for. Anthropic left it in.
  • HealthBench Professional — 59.8%. Opus 5 trails both Mythos 5 (66.0%) and GPT-5.6 Sol (60.5%). Clinical-adjacent work is not this model's strong suit.
  • Legal Agent Benchmark — 11.7% vs Fable 5's 13.3%, though both are far ahead of GPT-5.6 Sol's 2.5%. The absolute numbers here are low enough that the benchmark is mostly measuring how far there is left to go.
  • Humanity's Last Exam without tools (56.3% vs 56.5%) and FrontierCode Main (53.4% vs 53.5%) are ties inside the noise, not losses in any meaningful sense — but they do puncture the idea that Opus 5 simply dominates Fable 5.

The pattern: Opus 5 wins where the task is long-horizon, agentic, and requires the model to keep its own state straight across many steps. It draws or loses where the task is a single hard question answered in one pass. That is consistent with Anthropic's own framing of the release as a step change in "deep reasoning, agentic and long-horizon tasks, and test-time compute scaling" rather than a broad capability jump.

Anthropic also reports gains outside the public benchmarks: Opus 5 beats Opus 4.8 on every one of its life-sciences evaluations, with the largest deltas on organic chemistry (+10.2 percentage points on inferring molecular structures from spectroscopy data) and protein variant-effect prediction (+7.7 points).

The Effort Ladder Is the Real Story

Every headline number above is the top of a curve, and the curve is the thing worth understanding. Opus 5 exposes five effort levels, and Anthropic published the full ladder for each benchmark as score against cost per task.

Claude Opus 5 agentic coding performance on Frontier-Bench v0.1 by effort level, score versus cost per attempt, compared with Fable 5, Opus 4.8 and GPT-5.6 Sol

Read the Frontier-Bench chart and three things fall out.

First, the whole Opus 5 curve sits above and to the left of Fable 5's. At roughly $8.60 per attempt Opus 5 scores about 35%; Fable 5 needs around $27 to reach 33.7%. That is the "half the price" claim rendered as geometry rather than marketing.

Second, the curve bends. Going from low to medium buys about nine points for roughly $3. Going from xhigh to max buys nothing — the peak on this benchmark is at xhigh (≈44%), and max comes in slightly lower at the officially reported 43.3% while costing more. Anthropic's own documentation says as much: max "can be prone to overthinking." Treating max as strictly better is the most expensive mistake available here.

Third, GPT-5.6 Sol is genuinely cheaper at the low end. Its curve runs below Opus 5's on cost for the first several points on this benchmark and on the Artificial Analysis Coding Agent Index the two are effectively tied at the top. Opus 5 wins the peak; it does not own every point on the frontier.

The practical consequence is that carrying over effort defaults from Opus 4.8 is almost certainly wrong. Anthropic's guidance is to start at xhigh for coding and agentic work and high elsewhere, then sweep downward — because low and medium on Opus 5 outperform the same settings on prior Opus models by enough that a lot of production traffic can move down a rung without losing quality.

Knowledge Work: GDPval-AA and the Cost Frontier

Coding benchmarks get the attention, but the GDPval-AA v2 result is the one that should interest anyone doing document-heavy professional work. GDPval-AA is Artificial Analysis's evaluation built on OpenAI's GDPval dataset: real tasks drawn from 44 occupations, scored by blind pairwise comparison with an LLM judge and aggregated into an Elo.

Claude Opus 5 real-world knowledge task performance on GDPval-AA v2 by effort level, Elo score versus cost to run the full benchmark, compared with Fable 5, Opus 4.8 and GPT-5.6 Sol

Opus 5 reaches 1861 Elo, 114 points clear of Fable 5 and 268 clear of Opus 4.8. More usefully: it passes Fable 5's best score at roughly a third of Fable 5's cost to run the full benchmark. The one honest wrinkle is at the cheap end — GPT-5.6 Sol's curve starts higher, and only crosses below Opus 5's around the $300 mark. If your knowledge-work budget per task is very small, Opus 5 is not automatically the pick.

The same shape repeats on computer use. On OSWorld 2.0, Opus 5 tops out at 70.6% against Fable 5's 66.1%, and clears Fable 5's ceiling at just over a third of the cost.

What Changed Under the Hood

Opus 5 keeps the Opus 4.7/4.8 request surface almost entirely intact — budget_tokens is still gone, sampling parameters are still rejected, assistant prefills still return a 400. Two things genuinely change, and both can break working code.

Thinking is on by default. On Opus 4.8, a request that omitted the thinking field ran without thinking. On Opus 5, the same request thinks. The wire value did not change — thinking: {"type": "adaptive"} is still valid and equivalent to the default — but the default did. Because max_tokens caps thinking plus response text, a workload that sized max_tokens tightly around its answer on Opus 4.8 can now truncate mid-response. Revisit max_tokens on every route that never set a thinking field.

Disabling thinking is capped at high effort. Sending thinking: {"type": "disabled"} together with xhigh or max returns a 400. The check runs per request, so a later call that raises effort while thinking is still disabled fails even though earlier calls in the same conversation succeeded.

Anthropic also recommends against disabling thinking at all on this model, and documents two specific failure modes when you do. The model can occasionally write a tool call into its visible text instead of emitting a tool_use block — the turn completes normally, no error is raised, and the call simply never runs, which in an agentic loop leaves phantom text in the history that skews later turns. It can also leak `<thinking>` tags into the visible response. Counterintuitively, an instruction telling the model not to reason makes the tag leakage worse, not better. The recommended fix for both is the same: leave thinking on and use a lower effort level to control cost.

Three additions ship alongside:

  • Mid-conversation tool changes (beta header mid-conversation-tool-changes-2026-07-01) let you add or remove tools between turns without invalidating the prompt cache. Previously the tool list was fixed for a conversation's lifetime and any edit re-billed the entire prefix.
  • Automatic fallbacks gain a "default" mode (server-side-fallback-2026-07-01) that routes safety-classifier refusals to Anthropic's recommended fallback by refusal category, rather than a model list you maintain.
  • The prompt-cache minimum drops to 512 tokens, half of Opus 4.8's. Prompts you had written off as uncacheable may now create cache entries with no code change at all.

Pricing and Cost Economics

Opus 5 is $5 per million input tokens and $25 per million output tokens — identical to Opus 4.8 and exactly half of Fable 5's $10 / $50. Cache reads run $0.50 per million (a 90% discount) and cache writes $6.25. Fast mode doubles the base rate to $10 / $50 for roughly 2.5× output speed, and is Claude API only — it is not on Bedrock, Google Cloud, or Foundry.

Three cost levers matter more than the sticker price:

  1. Effort is the dominant term. On Frontier-Bench the spread from low to max is roughly 3× the cost per attempt. Moving a route from xhigh to medium where your evals allow it saves more than any prompt optimization will.
  2. Verbosity is a real bill. Opus 5's default responses and written deliverables run longer than prior Opus models', and lowering effort reduces thinking without reliably shortening visible output. A short explicit conciseness instruction is the lever, and it is worth writing.
  3. Delete your verification scaffolding. Opus 5 verifies its own work unprompted. Instructions like "include a final verification step" or "use a subagent to verify," carried over from earlier models, now cause over-verification. Removing them cuts tokens with no measured loss in quality — one of the rare optimizations that is a pure deletion.

Reported customer numbers back the efficiency story. A legal-tech customer saw comparable results with 26% fewer tokens than Opus 4.8 at max reasoning. A trading-benchmark customer reported roughly one-seventh the reasoning tokens and under half the latency. A financial-modeling customer averaged nine percentage points more accuracy with a third fewer turns and tool calls, and 60% less wall-clock time.

Alignment and Safety

Anthropic's automated behavioral audit scores Opus 5 at 2.30 on overall misaligned behavior, the lowest of its recent models — below Opus 4.8 (2.85), Mythos 5 (2.81) and Sonnet 5 (3.35), where lower is better on a 1–10 scale.

Automated behavioral audit results showing Claude Opus 5 scoring 2.30 on misaligned behavior, lower than Opus 4.8 at 2.85, Mythos 5 at 2.81 and Sonnet 5 at 3.35

On cybersecurity, Anthropic reports a split that is worth understanding rather than skimming: Opus 5 comes close to Mythos 5 at finding vulnerabilities, but remains substantially behind at exploiting them — turning a discovered vulnerability into a working attack. On the OSS-Fuzz evaluation the two models identify vulnerabilities at similar rates while Opus 5's exploit-development score lags far behind. Anthropic says it deliberately does not train Opus 5 on cyber tasks; the finding ability came along with general capability.

The practical upshot for developers is the classifier behavior. Opus 5's cyber safeguards allow finding vulnerabilities in source code but block binary-based vulnerability scanning, penetration testing, and exploit generation. Anthropic expects these classifiers to fire roughly 85% less often than Fable 5's — a direct response to the false-positive complaints that followed the Fable 5 launch, which we covered in our Fable 5 developer reception roundup. In Claude.ai, Claude Code and Claude Cowork, flagged requests fall back to Opus 4.8 by default; on the API, fallbacks are opt-in.

How to Use Claude Opus 5

If you are on Opus 4.8, the migration is a model-ID swap plus prompt re-tuning. The full checklist:

  1. Change the model string to claude-opus-5.
  2. Audit every route that disables thinking. Either enable thinking, or drop effort to high or below.
  3. Raise max_tokens on routes that never set a thinking field — they now think, and the cap covers both.
  4. Re-run an effort sweep on your own evals. Start xhigh for coding and agentic work, high elsewhere, then push downward.
  5. Delete verification instructions from prompts and verification steps from your harness.
  6. Add a conciseness instruction for user-facing routes, and a length instruction for Claude-authored files.
  7. Cap subagent delegation if your harness supports it — Opus 5 delegates more readily than Opus 4.8, which is the opposite of the problem 4.8 had.
  8. Re-check short prompts you had written off as uncacheable against the new 512-token minimum.

Two prompting notes that surprise people. Opus 5 follows severity filters literally, so a code-review harness that says "only report high-severity issues" will report less — ask it to report everything with a confidence and severity tag and filter downstream. And on vision work, giving the model tools to crop, inspect and visually verify its own output is a more cost-effective lever than raising thinking.

Who Should Use Claude Opus 5?

Use it if you run long-horizon agents, multi-file refactors, code review at volume, computer-use automation, or document-heavy enterprise workflows. It is also the obvious choice for anyone who wanted Fable 5's capability but could not accept its 30-day data-retention requirement or its aggressive cyber classifiers.

Consider alternatives if your workload is single-turn coding evaluated the way DeepSWE measures it (GPT-5.6 Sol still leads there), clinical or health-professional content (Mythos 5 and even GPT-5.6 Sol score higher), or cost-critical high-volume work where Claude Sonnet 5 at $3 / $15 closes much of the gap. For a direct comparison with the OpenAI side of the table, see our GPT-5.6 Sol guide.

What Claude Opus 5 Means for AI Slide Generation

A model that scores 1861 on real occupational knowledge work and holds instruction-following across a 1M-token window is, quietly, a document-to-deck model. Anthropic's own release notes list "office and document tasks" among the largest capability gains, specifically calling out well-structured slide decks and multi-sheet spreadsheets with non-trivial formulas — and one early-access customer building presentation software reported that the biggest gains showed up on "longer-horizon work: building a full deck, then revising it," with better visual understanding, cleaner formatting and fewer slide issues.

That maps onto the actual bottleneck in AI slide generation. Turning a 60-page report into a deck is not one task; it is three with different economics. Extraction is high-volume and low-ambiguity. Outlining — deciding the argument, the slide count, and what gets cut — is low-volume and almost pure judgment. Rendering and verification is a visual loop. Running all three at the same effort level is how document-to-PPT pipelines end up either expensive or shallow.

Diagram mapping Claude Opus 5 effort levels onto the three stages of a document-to-deck pipeline, with the published Frontier-Bench cost curve underneath

The effort ladder makes that split addressable inside one model. Extraction runs at low or medium, where Opus 5 now beats prior Opus models at their own settings. Outlining runs at high or xhigh, where the long-horizon reasoning gains actually show up and where a bad structural call costs far more than the tokens saved. Rendering and layout verification benefits less from more thinking than from tools — Anthropic is explicit that Opus 5's vision performance is strongest when it can crop, inspect and re-check its own output, which is exactly the loop a slide renderer needs to catch an overflowing text box or a chart that has drifted off the grid. Our HTML vs image slide generation guide covers why that verification step is where the two rendering approaches diverge most.

The one behavior to prompt around is verbosity. Opus 5 writes longer deliverables by default, and on a slide deck "longer" means bullet points that overflow the frame. An explicit length instruction per slide role — cover, section, content, closing — is now part of the prompt, not an optional refinement. The same discipline applies to the outline stage covered in our PDF-to-PowerPoint conversion guide: a model that expands scope unprompted will happily add three slides nobody asked for.

At Tosea.ai this is the orchestration layer we sit on. A raw frontier model gives you an excellent single answer; a document-to-deck product has to decide which model, at which effort, for which stage, and then hold the result to a slide structure that survives contact with a real audience. Opus 5 does not remove that layer — it makes the cost curve underneath it much more favorable, and gives the outline stage a genuinely better reasoner to work with. For the structural side of that problem, our board deck structure guide covers what the narrative stage is actually deciding.

Frequently Asked Questions

Is Claude Opus 5 better than Claude Fable 5? On most benchmarks in the official table, yes — and at half the price. But Fable 5 still edges it on Humanity's Last Exam without tools and the Legal Agent Benchmark, and Mythos 5 leads on health and long-horizon biology work. Anthropic itself says Opus 5 "comes close to" Fable 5's frontier intelligence rather than surpassing it.

Does Claude Opus 5 cost more than Opus 4.8? No. Identical pricing at $5 / $25 per million tokens. Fast mode is $10 / $50.

What is the biggest breaking change from Opus 4.8? Thinking is on by default, and thinking: {"type": "disabled"} now returns a 400 at xhigh or max effort. Both can break code that worked on Opus 4.8.

Which effort level should I use? Start at xhigh for coding and agentic work and high for everything else, then sweep downward on your own evaluations. max is not reliably better — on Frontier-Bench it scores slightly below xhigh while costing more.

Is Opus 5 on LMArena yet? Not at the time of writing. It launched July 24, 2026, and arena ratings typically need one to two weeks of votes to stabilize. Artificial Analysis has already placed it first on its Intelligence Index at 61.

Can I use it for security research? Partly. Finding vulnerabilities in source code is allowed; binary-based vulnerability scanning, penetration testing and exploit generation are blocked. Anthropic's Cyber Verification Program provides access to a less restricted version for enterprises and researchers.

The Bottom Line

Claude Opus 5 is the rare launch where the pricing page is as interesting as the benchmark chart. Anthropic did not raise the price, published cost curves instead of single scores, left four losing rows in its own comparison table, and shipped a model that its independent evaluator ranks first overall while flagging it as unusually verbose.

For agentic and long-horizon work the case is straightforward — the Frontier-Bench and AutomationBench margins are large enough that no amount of harness tuning closes them. For single-pass coding, GPT-5.6 Sol still has a defensible claim, and for clinical work Opus 5 is not the answer. The migration itself is small: one model string, one thinking audit, one effort sweep, and a round of deleting the verification instructions you spent last year writing.

Sources

Continue Reading

All Insights