How to Use Claude Sonnet 5.5: Complete Guide to Benchmarks, Cost per Task and Migration
Claude Sonnet 5.5 launched September 28, 2026 at Sonnet 5 pricing. Full official benchmarks, independent Artificial Analysis results, real cost per task by effort, and API migration changes.
What Is Claude Sonnet 5.5?
Anthropic released Claude Sonnet 5.5 on September 28, 2026, six days after Claude Opus 5.5. It is the second model in the Claude 5.5 family and replaces Claude Sonnet 5 as Anthropic's mid-tier model. Claude Haiku 5.5 is still listed as coming "in the coming weeks."
The launch pitch has three parts. Sonnet 5.5 is a large upgrade over Sonnet 5, most visibly in agentic coding, where its Terminal-Bench 4.0 score rises from 10.3% to 70.6%. It generates output more than 30% faster. And it keeps Sonnet 5's list price of $2 per million input tokens and $10 per million output tokens, which Anthropic says works out to up to 30% lower cost per task because the model needs fewer tokens and tool calls for the same work.
Anthropic also draws a clear line between its two new models. Opus 5.5 is "built for complex work requiring careful judgment," while Sonnet 5.5 is "strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets." That last phrase is the reason we read this launch closely: turning documents into slides is our core workload.
This guide covers the full official benchmark table, the independent results Artificial Analysis published on launch day, what the per-effort cost curves actually show, the API changes you need before switching, and how to choose between Sonnet 5.5 and Opus 5.5 for a given job.
A disclosure before we start: Tosea.ai builds an AI presentation tool that routes work across several model vendors, including Anthropic and OpenAI. We have no commercial relationship with either company beyond paying for API usage.
Claude Sonnet 5.5 at a glance
| Spec | Claude Sonnet 5.5 |
|---|---|
| Release date | September 28, 2026 |
| API model ID | claude-sonnet-5-5 |
| Context window | 1M tokens, billed at the standard rate across the full window |
| Max output | 128K tokens (300K on the Batch API with a beta header) |
| Price (input / output) | $2 / $10 per million tokens |
| Cache reads / 5-minute cache writes | $0.20 / $2.50 per million tokens |
| Batch API (input / output) | $1 / $5 per million tokens |
| Effort levels | low, medium, high, xhigh, max |
| Default effort | high on the Claude Platform; medium in Claude Code and the Claude apps |
| Thinking | Adaptive by default; between_tools is the lowest setting |
| Reliable knowledge cutoff | June 2026 |
| Availability | Claude apps, Claude Code, Claude Platform, AWS, Google Cloud, Microsoft Azure |
Claude Sonnet 5.5 Benchmarks: The Full Official Table
Anthropic compares Sonnet 5.5 with Sonnet 5, Opus 5.5 and OpenAI's GPT-6 Sol, which carries the same $2/$10 list price. Here is the complete table, including the rows where Sonnet 5.5 does not come first.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 70.6% | 10.3% | 66.4% | — |
| FrontierCode 1.1 Main (mergeable code) | 46.2% (52.1% at xhigh) | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| GDPval-AA v2.1 (knowledge work, Elo) | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 (long-horizon knowledge work, Elo) | 1811 | 1359 | 1822 | 1483 |
| Humanity's Last Exam (with tools) | 64.5% | 54.9% | 67.7% | — |
| OSWorld 2.1 (computer use, partial credit) | 80.1% | 57.0% | 81.8% | — |
| Chartography (chart reading, no tools) | 61.6% | 15.6% | 64.4% | 53.6% |
Source: Anthropic's launch page. The headline Sonnet 5.5 numbers are max-effort results; the cost charts on the same page confirm this for the four coding and knowledge-work rows. Opus 5.5's Terminal-Bench 4.0 score is its best result, which it reached at xhigh. Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment that had a structured-outputs bug, since fixed. GPT-6 Sol's AA-Briefcase, GDPval-AA and Chartography scores may predate an OpenAI fix for an image-understanding bug.
Two footnotes change how the table reads.
The first is the FrontierCode row. Sonnet 5.5 scores lower at max effort (46.2%) than at xhigh (52.1%). FrontierCode checks whether a change could be merged without human edits, and it penalizes out-of-scope changes even when they are useful. At max effort, Sonnet 5.5 more often ran Claude Code's code-review skill, which splits the review across many subagents, and in the cases Cognition examined this led to timeouts or edits beyond the task. More effort is not always better.
The second is the Chartography row, which uses the no-tools setting. The Opus 5.5 launch page reported 89.0% for Opus 5.5 with tools. The two numbers measure different things, so do not compare them across launch pages.
Where Sonnet 5.5 Wins, Ties, and Trails
Against Sonnet 5: a generational jump. Every row improves, several by very large margins. Terminal-Bench 4.0 goes from 10.3% to 70.6%, Chartography from 15.6% to 61.6%, OSWorld from 57.0% to 80.1%, and GDPval-AA gains 395 Elo points. If you run Sonnet 5 in production today, no row on this table favors staying.
Against Opus 5.5: close, but not equal. Sonnet 5.5 leads Opus 5.5 on exactly one row, Terminal-Bench 4.0, by 70.6% to 66.4%. It trails narrowly on GDPval-AA (2 Elo points), OSWorld (1.7 points), CursorBench (2.3 points) and Chartography (2.8 points), and by a bit more on AA-Briefcase (11 Elo points) and Humanity's Last Exam (3.2 points). FrontierCode shows the widest gap: 54.4% for Opus 5.5 against Sonnet 5.5's best of 52.1% at xhigh and 46.2% at max. Anthropic adds the qualifier that matters most: "in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment."
Against GPT-6 Sol: favorable at the same list price. On the four rows where Anthropic reports GPT-6 Sol, Sonnet 5.5 leads on AA-Briefcase (1811 vs 1483), GDPval-AA (1844 vs 1487) and Chartography (61.6% vs 53.6%). On FrontierCode it trails at max (46.2% vs 49.3%) but leads at xhigh (52.1%). The per-task cost story is more complicated than the scores, as we show below. Our GPT-6 Sol and Luna guide covers OpenAI's side of this comparison.
What is missing. Anthropic reports no GPT-6 Sol figures for Terminal-Bench 4.0 or CursorBench, because neither has been published, and its cost charts substitute GPT-5.6 Sol on those two. There is no GPT-6 Astra or Claude Fable 5.1 column, so this table alone cannot place Sonnet 5.5 against the top closed models. Artificial Analysis fills part of that gap.
Independent Results: Artificial Analysis
Artificial Analysis evaluated Sonnet 5.5 at all five effort levels on launch day. On its Intelligence Index v4.3.2, which combines ten evaluations, Sonnet 5.5 at max effort scores 56. That is second only to Opus 5.5 (max) at 58, ahead of Claude Fable 5.1 and GPT-6 Astra (max) at 53, and 18 points above Sonnet 5 (max) at 38.

The sub-scores show where the 56 comes from and where it does not:
| Evaluation (Artificial Analysis harness, max effort) | Sonnet 5.5 | Opus 5.5 | Sonnet 5 |
|---|---|---|---|
| Intelligence Index v4.3.2 | 56 | 58 | 38 |
| Terminal-Bench 4.0 | 63.6% | 59.6% | 14.1% |
| AutomationBench-AA | 71% | 70% | 37% |
| GDPval-AA v2.1 (Elo) | 1844 | 1846 | 1449 |
| AA-Briefcase v1.1 (Elo) | 1811 | 1822 | 1359 |
| GDP.pdf (professional document reasoning) | 26% | 26% | 13% |
| AA-LCR (long-context reasoning) | 83% | 85% | 82% |
| SciCode | 61% | 67% | 54% |
| Humanity's Last Exam | 55% | 61% | 41% |
| AA-Omniscience accuracy | 54% | 66% | 40% |
| AA-Omniscience hallucination rate (lower is better) | 47% | 59% | 39% |
Three findings stand out.
The agentic lead holds up independently. Under Artificial Analysis's own harness, Sonnet 5.5 still tops Terminal-Bench 4.0 at 63.6%, ahead of Opus 5.5 (59.6%) and GPT-6 Astra (59.1%). It also leads AutomationBench-AA, a test of agentic SaaS workflows, at 71% against 70% for Opus 5.5. This is the reverse of what happened with the Opus 5.5 launch, where Artificial Analysis's harness shrank Anthropic's Terminal-Bench lead to a tie.

Knowledge and science still favor Opus. Sonnet 5.5 is 6 points behind Opus 5.5 on SciCode and Humanity's Last Exam, and 12 points behind on AA-Omniscience factual accuracy. On Terminal-Bench-Science, a benchmark of scientific research workflows in the terminal, it scores 53.3%, behind GPT-6 Astra (63.3%) and Opus 5.5 (59.0%). A smaller model carries less stored knowledge, and these rows show it.
Hallucination is a mixed picture. Sonnet 5.5's hallucination rate (47%) is lower than Opus 5.5's (59%), meaning it more often declines when it does not know. It is higher than Sonnet 5's 39%, though, and Sonnet 5.5 also answers far more questions correctly (54% against 40%). If your workflow penalizes confident errors more than it rewards coverage, check this row before migrating.
The Token Bill: Where "30% Cheaper" Holds and Where It Doesn't
Anthropic's cost claim and Artificial Analysis's cost finding look contradictory until you separate effort levels.
Artificial Analysis measured about 193K output tokens per Intelligence Index task for Sonnet 5.5 at max effort. That is the highest figure it has recorded for any model: roughly 60% more than Opus 5.5 (max) or Sonnet 5 (max), and about seven times GPT-6 Astra (max). Artificial Analysis puts the resulting cost at $7.60 per task, about 50% more than Sonnet 5. On Artificial Analysis's cost chart, Sonnet 5.5 at max sits slightly to the right of Opus 5.5 at max, meaning it costs a little more per task for a score 2 points lower.

The lower settings tell a different story. Sonnet 5.5 used 14K tokens per task at low, 19K at medium, 34K at high and 74K at xhigh. Anthropic's own charts, which plot score against cost per task at every effort level, show where the "up to 30%" and "a tenth of the cost" claims come from:
- Terminal-Bench 4.0: at medium, Sonnet 5.5 scores 28.8% for $0.83 per attempt, well above Sonnet 5's best of 10.3% at $11.62.
- FrontierCode: at high, Sonnet 5.5 scores 49.4% for $0.42 per task, 10 points above Sonnet 5 at high (39.4% for $6.10), roughly one-fifteenth of the cost. That also matches GPT-6 Sol's best score (49.3%, at $2.07) for about a fifth of the cost.
- CursorBench 4.0: at low, Sonnet 5.5 scores 35.8% for $0.50, above Sonnet 5's best of 34.1% at $7.17.
- AA-Briefcase: at medium, Sonnet 5.5 reaches 1461 Elo for $1.64, above Sonnet 5's best of 1359 at $14.43.
Against Sonnet 5, then, the savings are real and large at every setting below max. Against Opus 5.5, the same data is less flattering, and it is the part of the launch that deserves the most attention.
On Terminal-Bench 4.0, Opus 5.5 at high scores 64.2% for $3.88 per attempt. Sonnet 5.5 at xhigh scores 61.5% for $5.30. Opus is both better and cheaper in that range. Sonnet 5.5 only passes Opus's best score (66.4%) at max, where it spends $12.54 per attempt against $7.35 for Opus at xhigh, about 70% more for 4.2 extra points.
On AA-Briefcase, Sonnet 5.5 at high (1634 Elo, $3.95) and Opus 5.5 at medium (1642 Elo, $4.40) are effectively the same point. At the top end, Sonnet 5.5 at max costs $29.19 per task for 1811 Elo, while Opus 5.5 at max costs $21.05 for 1822.
On FrontierCode, Sonnet 5.5 at high (49.4%, $0.42) edges Opus 5.5 at low (47.3%, $0.40). Above that price, Opus 5.5 at medium reaches 54.6% for $0.80, the best score on the chart, at half the cost of Sonnet 5.5 at xhigh.
The pattern is consistent. Sonnet 5.5 is the cost-efficient choice at low, medium and high effort. At xhigh and max, Opus 5.5 usually delivers the same or better score for the same or less money. Artificial Analysis reaches a similar conclusion: Sonnet 5.5 sits off its intelligence-versus-cost Pareto frontier, behind Opus 5.5 at high effort levels and behind GPT-6 configurations at low ones, with its high setting the most competitive, just behind GPT-6 Sol at about the same cost per task.
Claude Sonnet 5.5 Pricing
| Price per 1M tokens | Sonnet 5.5 | Opus 5.5 | Sonnet 5 |
|---|---|---|---|
| Input | $2 | $4 | $2 |
| Output | $10 | $20 | $10 |
| Cache reads | $0.20 | $0.20 | $0.20 |
| Cache writes (5 minutes) | $2.50 | $5 | $2.50 |
| Cache writes (1 hour) | $4 | $8 | $4 |
| Batch input / output | $1 / $5 | $2 / $10 | $1 / $5 |
A few details from Anthropic's pricing documentation are easy to miss:
- Sonnet 5's price did not go up. Sonnet 5 launched with $2/$10 labeled as introductory pricing through August 31, 2026, with a scheduled rise to $3/$15. Anthropic has since made $2/$10 the standard price, so Sonnet 5.5 is "the same price" as Sonnet 5 in the literal sense.
- Long context costs nothing extra. A 900K-token request is billed at the same per-token rate as a 9K-token request, and caching and batch discounts still apply.
- Caching gets easier. The minimum cacheable prompt drops from 1,024 tokens on Sonnet 5 to 512 tokens.
- Cache reads are the same price as on Opus 5.5. Both models charge $0.20 per million cache-read tokens, so for agent loops dominated by cached context, the per-token gap between them is smaller than the list prices suggest.
The practical implication follows from the previous section. Per-token pricing is half of Opus 5.5, but per-task cost depends mostly on how many tokens the model spends. Measure cost per completed task on your own workload before assuming the halved list price carries through.
How to Choose an Effort Level
Sonnet 5.5 has five effort levels. They are recalibrated from Sonnet 5, so the same setting does not produce the same amount of thinking. Anthropic's migration guide recommends starting at high for most work, medium for well-specified agentic coding and multistep tool use, and medium or low for chat and latency-sensitive tasks.
| Effort | Use it for | What the data suggests |
|---|---|---|
| low | Classification, extraction, routing, short rewrites | About 14K tokens per task; already beats Sonnet 5's best on CursorBench |
| medium | Well-specified coding, drafting, everyday agent steps | About 19K tokens per task; beats Sonnet 5's best on Terminal-Bench and AA-Briefcase |
| high (API default) | Multi-file changes, analytical documents, reports | About 34K tokens per task; the setting Artificial Analysis rates most competitive |
| xhigh | Long agentic sessions, hard debugging | Best FrontierCode score; at this cost, test Opus 5.5 at high as well |
| max | Benchmarks and the hardest single tasks | About 193K tokens per task; Opus 5.5 is often cheaper for the same score |
Our default: run Sonnet 5.5 at medium or high, measure where it fails, and escalate only those steps. If a step needs xhigh or max, run the same step on Opus 5.5 at high or medium and compare cost per completed task. That comparison will often favor Opus.
What Changed in the API
Sonnet 5.5 is not a drop-in model swap. Anthropic's migration guide lists several changes for code coming from Sonnet 5, and some return errors rather than degrading quietly.
- Thinking cannot be disabled.
thinking: {"type": "disabled"}returns a 400 error. The lowest setting is the newbetween_toolsmode, which skips up-front thinking but still returns progress updates between tool calls. It works only at low, medium and high effort; at xhigh or max it returns a 400. - Forced tool use is gone. A
tool_choiceof typeanyortoolreturns a 400 error, including on the token-counting endpoint. Useautowithstrict: trueon the tools that need schema-exact input, and say in the prompt when each tool should be called. - Conversations must be append-only. Each thinking block is signed over the conversation before it. For accounts created on or after August 31, 2026, replaying a block after editing earlier history returns a 400. Agent frameworks that rewrite or trim prior turns need testing.
- Thinking blocks do not cross model families. Sonnet 5.5 reads thinking blocks from Sonnet 5 and older models, but not from Opus 5, Opus 5.5, Fable or Mythos models. The API drops unreadable blocks silently, so a mid-conversation handoff from Opus 5.5 loses that reasoning.
- Computer use moves to a toolset. On the Claude API and Google Cloud, computer use requires
computer_toolset_20260801; the oldercomputer_20251124returns a 400 error there. - Text between tool calls moves into thinking blocks. Longer notes the model writes between tool calls now arrive as progress-update thinking blocks, empty at the default display setting. If your interface shows those notes to users, set
displayto"summarized"or use the updates display beta. - More refusal categories. Declines return
stop_reason: "refusal"with categories includingcyber,bio,frontier_llmandreasoning_extraction. Server-side fallback, which is in beta, retries cyber and frontier-LLM declines on Sonnet 5.
Anthropic also ships a migration helper: running /claude-api migrate in Claude Code applies the model ID swap and the breaking parameter changes across a codebase, then produces a checklist of items to verify manually.
Safety and Safeguards
Anthropic says Sonnet 5.5 "doesn't advance the frontier of our models' capabilities," so its alignment assessment focused on risks that apply at any capability level, such as acting against users' interests, misleading users, and cooperating with high-stakes misuse. On an automated behavioral audit of roughly 1,850 scenarios, Sonnet 5.5 improves on or matches Sonnet 5 on most measures. In the newer containment evaluations it comes close to Opus 5.5 in how rarely it tries to escape its sandbox. Across the full audit, Opus 5.5 still performs slightly better.
The safeguards are stricter than for any previous Sonnet. Because its cyber capabilities are comparable to Opus 5's, Sonnet 5.5 is the first Sonnet model to launch with the cyber safeguards Anthropic uses on its top models. Routine bug finding and fixing is unaffected, but higher-risk security tasks visibly fall back to Sonnet 5. It is also the first Sonnet with classifiers that block reasoning extraction, part of Anthropic's response to distillation attacks. Artificial Analysis saw fallbacks on about 0.1% of Intelligence Index tasks, mostly in Terminal-Bench 4.0.
The Next Web pointed out a tension in the messaging: a model described as not advancing the frontier ships with frontier-style cyber restrictions. Both statements can be true, since "frontier" here refers to Anthropic's own best models rather than to any single capability, but it is a fair reminder that launch language compresses a lot of nuance.
Who Should Use Claude Sonnet 5.5?
- Teams on Sonnet 5. Upgrade after handling the API changes above. Every official benchmark improves, and at medium or high effort the cost per task falls sharply.
- Teams that use Opus 5.5 at low or medium for well-scoped work. Test Sonnet 5.5 at high. On AA-Briefcase and FrontierCode it reaches similar scores at similar or lower cost, and it runs faster.
- Terminal-heavy coding agents. Sonnet 5.5 has the highest Terminal-Bench 4.0 score in both Anthropic's table and Artificial Analysis's leaderboard, but that lead comes at max effort. Compare it with Opus 5.5 at high on your own tasks.
- Knowledge-heavy or scientific work. Opus 5.5 or GPT-6 Astra remain better choices on factual recall, SciCode and Terminal-Bench-Science.
- Maximum volume at minimum cost. GPT-6 Sol at low or medium effort is cheaper per task on several of these charts, and Haiku 5.5 is still to come. See our GPT-6 Sol and Luna guide.
What Claude Sonnet 5.5 Means for AI Slide Generation
Anthropic markets Sonnet 5.5 explicitly for "polished documents, slides, and spreadsheets," and says early testers found it "can follow slide templates to create decks that require minimal editing." In one internal test, Anthropic gave the model a public company's quarterly earnings materials and call transcripts, plus a slide template, and asked for a 10-slide operating review. Two experts judged the first draft ready to send. That is a vendor-selected example, but it is exactly the document-to-deck task we see most often; our earnings call presentation guide covers the structure such a deck needs.
Three parts of this release matter for an AI presentation tool.
Reading the source. Chartography at 61.6% without tools, up from 15.6% on Sonnet 5, is one of the largest jumps in the table. In a PDF-to-PowerPoint workflow, charts in the source have to be read accurately before they can be rebuilt as editable slide charts, and Sonnet 5 was weak at this. The 1M-token context at standard pricing means a full annual report or thesis fits in one request. GDP.pdf, Artificial Analysis's professional document-reasoning test, puts Sonnet 5.5 level with Opus 5.5 at 26%, which also shows that even the best models still get many document questions wrong.
Deciding the slide structure. Choosing the storyline and which evidence supports each point is where Opus 5.5 still earns its price. Anthropic's own "sustained judgment" caveat applies directly to outline generation. At similar cost per task, Sonnet 5.5 at high and Opus 5.5 at medium score about the same on AA-Briefcase, which grades presentation quality alongside analysis, so this is a measurable choice rather than a guess.
Writing and revising slides. Per-slide content, template-following, and short revisions are well-scoped tasks, the category where Sonnet 5.5 is strongest and where low and medium effort are enough. A 20-slide deck involves dozens of such calls, so a faster model with half the per-token price changes the economics of iterative editing.
This split between a stronger model for structure and a faster one for the per-slide work is how Tosea.ai approaches the document-to-deck pipeline, choosing the model and effort level per stage rather than running the most expensive setting everywhere. For keeping numbers traceable to the source, see our zero-hallucination slides guide, and for the full presentation workflow, our PDF to PowerPoint guide.
Frequently Asked Questions
When was Claude Sonnet 5.5 released? September 28, 2026, on the Claude apps, Claude Code, the Claude Platform, AWS, Google Cloud and Microsoft Azure.
How much does Claude Sonnet 5.5 cost? $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. Cache reads cost $0.20, five-minute cache writes $2.50, and the Batch API halves the base rates to $1/$5.
Is Claude Sonnet 5.5 better than Claude Opus 5.5? On one official benchmark, Terminal-Bench 4.0, yes. On the other seven, Opus 5.5 scores higher, mostly by 1 to 3 points, and Artificial Analysis ranks Opus 5.5 at 58 against 56 for Sonnet 5.5. Opus 5.5 is also stronger on factual knowledge and scientific reasoning.
Is Sonnet 5.5 cheaper than Opus 5.5 per task? At low, medium and high effort, usually yes. At xhigh and max, Sonnet 5.5 can use so many more tokens that Opus 5.5 reaches the same score for the same or less money.
Can I turn off thinking in Sonnet 5.5?
Not completely. The disabled setting returns an error. Use between_tools, which removes up-front thinking, at low, medium or high effort.
What is the context window? 1M tokens, with up to 128K output tokens on the standard Messages API.
Is Haiku 5.5 available? Not yet. Anthropic says it will join the Claude 5.5 family in the coming weeks.
The Bottom Line
Claude Sonnet 5.5 is a large upgrade over Sonnet 5 at the same list price, and close enough to Opus 5.5 on agentic coding and knowledge work that, for many tasks, choosing between them is a cost question rather than a quality one. The independent numbers back most of Anthropic's claims, including the Terminal-Bench lead. The cost claim needs a qualifier: the savings are real at low, medium and high effort, but at max effort Sonnet 5.5 is the most token-hungry model Artificial Analysis has measured and often costs more per task than Opus 5.5. The practical approach is to use Sonnet 5.5 at medium or high for well-scoped work and send the steps that need maximum effort to Opus 5.5.
Sources
- Introducing Claude Sonnet 5.5 — Anthropic, September 28, 2026
- Claude Sonnet 5.5 System Card — Anthropic
- Migrating to Claude Sonnet 5.5 — Claude Platform documentation
- Models overview and Pricing — Claude Platform documentation
- Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Index — Artificial Analysis, September 28, 2026
- Anthropic releases Sonnet 5.5, which it calls a significantly cheaper, faster work partner — TechCrunch
- Anthropic launches Claude Sonnet 5.5 with 30% cost reduction per-task due to faster speeds and fewer tool calls — VentureBeat
- Anthropic debuts Claude Sonnet 5.5 running 30% faster than the previous-generation AI model — SiliconANGLE
- Anthropic releases Claude Sonnet 5.5 with the cyber limits it reserved for its best models — The Next Web