Ranked #3 Coding — AI That Writes Production Code
Anthropic

Claude Sonnet 5.5

Claude Sonnet 5.5 is the coding model you can leave switched on all day: fast, half the token price of Opus 5.5, and good enough that most bug fixes, features, and refactors never need the flagship. Just don't turn the effort dial all the way up — at its highest setting it writes more than any model Artificial Analysis has measured, and some results get worse.

Updated September 29, 2026 Agentic CodingTerminal-Bench 4.0 64%Claude Code
9.8out of 10
Official Website
Best for

Claude Sonnet 5.5 is the coding model you can leave switched on all day: fast, half the token price of Opus 5.5, and good enough that most bug fixes, features, and refactors never need the flagship. Just don't turn the effort dial all the way up — at its highest setting it writes more than any model Artificial Analysis has measured, and some results get worse.

Why It Wins

Independent Artificial Analysis testing puts it at 64% on Terminal-Bench 4.0, slightly above Opus 5.5 and GPT-6 Astra (about 60% each). Anthropic reports 52.1% on FrontierCode at xhigh effort and 55.5% on CursorBench 4.0, within about two points of Opus 5.5. Vals AI ranks it second of 66 models on its index and first on Vibe Code Bench and Code Migration. Tokens cost $2/$10, and output streams 30%+ faster than Sonnet 5.

Watch out

At max effort it used about 193,000 output tokens per task in Artificial Analysis testing, the most they have measured, and its FrontierCode score drops from 52.1% at xhigh to 46.2% at max. Anthropic itself says Opus 5.5 remains clearly stronger at complex, open-ended work. Higher-risk security tasks fall back to Sonnet 5, and API users face five breaking changes when migrating.

01

What It Actually Is

Picture two carpenters. One is the master you call when a house needs a new staircase and nobody has drawn the plans. The other is the steady tradesperson who can hang forty doors in a week, every one of them square.

Most of a software team’s week is doors, not staircases: a failing test, a form that needs a new field, an API that changed its name. Claude Sonnet 5.5 is built for the doors.

Anthropic released it on September 28, 2026, as the second model in the Claude 5.5 family, a week after Opus 5.5. It keeps Sonnet 5’s price of $2 per million input tokens and $10 per million output, which is half of Opus 5.5. Anthropic says it runs more than 30% faster than Sonnet 5 and costs up to 30% less per task because it finishes in fewer steps.

The independent numbers

For coding agents, the test we trust most is Terminal-Bench 4.0: long, multi-step jobs in a real command-line environment, where the model has to install things, run them, read the errors, and try again.

Artificial Analysis ran it independently in September 2026. Sonnet 5.5 scored 64%, slightly ahead of Opus 5.5 and GPT-6 Astra at about 60% each. Anthropic’s own run is higher, at 70.6%, but vendor harnesses tend to flatter their own models, so the independent 64% is the number to remember. Either way, it is an enormous jump from Sonnet 5, which scored 10.3% in Anthropic’s table.

A second independent source agrees on the big picture. Vals AI places Sonnet 5.5 second of 66 models on its index at 69.22%, less than half a point behind Opus 5.5, while costing about two-thirds as much per test. It takes first place on two coding leaderboards: Vibe Code Bench (building small apps from a description) at 92.39% and Code Migration at 69.83%.

What Anthropic’s own tests add

Anthropic’s launch table fills in the details:

  • FrontierCode v1.1, which asks whether a code change could be merged without a human touching it: 52.1% at xhigh effort. Opus 5.5 scores 54.4% and GPT-6 Sol 49.3%.
  • CursorBench 4.0, built from real sessions in the Cursor editor: 55.5%, about two points behind Opus 5.5 and far ahead of Sonnet 5’s 34.1%.

The more useful part is the cost charts, which plot every model’s score against its cost per task at each effort level. Two findings stand out. At High effort, the API default, Sonnet 5.5 matches GPT-6 Sol’s best FrontierCode score for about a fifth of the cost per task. And at Medium effort, the default in Claude Code, it beats Sonnet 5’s best Terminal-Bench score for less than a tenth of the cost.

Early testers describe the same thing in plainer words. Lovable saw roughly half as many shell commands per task. Slack measured about 14% fewer output tokens than Sonnet 5. Unity, which only counts a task as done when the change actually works at runtime, says the majority of Sonnet 5.5’s work passed that check.

The effort dial, and why the top setting backfires

Like Opus 5.5, Sonnet 5.5 has an effort dial with five settings: low, medium, high, xhigh, and max. Higher settings let the model think longer and check its work more. You would expect max to be best. For Sonnet 5.5, it often isn’t.

Artificial Analysis measured the whole range on its Intelligence Index:

Effort Intelligence Index Cost per task
Low 36 $0.41
Medium 41 $0.59
High 47 $1.08
Xhigh 52 $2.74
Max 56 $7.60

Going from xhigh to max buys four more points for nearly three times the cost. At max, Sonnet 5.5 wrote about 193,000 output tokens per task. That is the most Artificial Analysis has ever measured: about 60% more than Opus 5.5 at max, and about seven times GPT-6 Astra.

Worse, more effort does not always mean better code. On FrontierCode, Sonnet 5.5 scores 52.1% at xhigh but 46.2% at max. Anthropic explains why in a footnote: at max, the model more often launched Claude Code’s code-review routine, which splits the review across many helper agents. In some cases that timed out or made extra edits beyond the task, and FrontierCode penalizes changes nobody asked for.

The lesson is simple: treat xhigh as the ceiling. If a task still fails at xhigh, the better next step is usually Opus 5.5, not max.

The honest catch

Opus is still the architect. Anthropic says it plainly: on several tests Sonnet 5.5 at max comes close to Opus 5.5, but in Anthropic’s own testing and that of outside testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment. Opus also keeps the lead on FrontierCode and CursorBench.

The Claude Code bill won’t halve. Coding agents re-read your project constantly, and those re-reads are billed as cache reads, which cost $0.20 per million tokens on both Sonnet 5.5 and Opus 5.5. You save on the new input and output, but a long cache-heavy session saves less than the headline “half price” suggests.

Security work is rerouted. Sonnet 5.5 is the first Sonnet to ship with the cyber safeguards Anthropic uses on its top models. Routine bug fixing is unaffected, but higher-risk security tasks visibly fall back to Sonnet 5 unless your team is approved through Anthropic’s Cyber Verification Program.

Migration takes a little work. Anthropic lists five breaking changes from Sonnet 5. Forced tool use now returns an error, non-default temperature settings are rejected, and text between tool calls comes back in a different format. None of this is hard to fix, but it is not a one-line model swap.

Where it fits

If the job is… Reach for Why
A scoped bug, feature, or migration Claude Sonnet 5.5 Fast and cheap at medium or high effort
Open-ended architecture or a problem nobody has scoped Claude Opus 5.5 Clearly stronger judgment, per Anthropic and testers
Long unattended terminal work at the lowest cost per task GPT-6 Astra Uses a fraction of the tokens at peak
Front-end polish, slides, small visual builds Claude Sonnet 5.5 A sharp eye for design and fast iteration

The verdict

Claude Sonnet 5.5 is the model most developers should leave switched on. It handles the doors, all forty of them, quickly and at half the flagship’s token price. Keep it at medium or high effort, use xhigh for the hard ones, and when the job turns out to be a staircase, call Opus.

02

Strengths and honest limitations

Key Strengths

  • Independent terminal results: Artificial Analysis measured 64% on Terminal-Bench 4.0 in September 2026, slightly above Opus 5.5 and GPT-6 Astra at about 60%. That is the independent test that decides most coding-agent comparisons, and Sonnet 5.5 did well on it at half Opus’s token price.
  • Cheap at sensible settings: Anthropic’s charts show Sonnet 5.5 at High effort, the API default, matching GPT-6 Sol’s best FrontierCode score for about a fifth of the cost per task. On Terminal-Bench, Medium effort beats Sonnet 5’s best score for less than a tenth of the cost.
  • Finishes in fewer steps: Early testers report the same pattern. Slack measured about 14% fewer output tokens than Sonnet 5, Lovable saw roughly half the shell runs per task, and a Balyasny test used about 121,000 tokens per answer where Sonnet 5 used 497,000.
  • Strong on migrations and quick builds: Vals AI ranks Sonnet 5.5 first on Vibe Code Bench (92.39%) and Code Migration (69.83%), and second overall on the Vals Index at about two-thirds of Opus 5.5’s cost per test.
  • A good eye for interfaces: Anthropic says Sonnet 5.5 has a sharp eye for design and can polish user interfaces and follow slide templates closely enough that little editing is needed. Output also arrives 30%+ faster than Sonnet 5, which makes quick back-and-forth editing pleasant.

Honest Limitations

  • Max effort is a trap: Artificial Analysis measured about 193,000 output tokens per task at max, roughly 60% more than Opus 5.5 at max and about seven times GPT-6 Astra. At that setting most of the price advantage over Opus disappears.
  • More effort, worse merges: On FrontierCode, which checks whether changes could be merged without human edits, Anthropic reports 52.1% at xhigh but 46.2% at max. At max, the model more often launched a code-review routine split across many helper agents, which sometimes timed out or edited beyond the task.
  • Not the architect: Anthropic says Opus 5.5 remains clearly stronger at complex, open-ended work that needs sustained judgment. Opus also leads FrontierCode (54.4%) and CursorBench (57.8%).
  • Security work is rerouted: Higher-risk cybersecurity tasks visibly fall back to Sonnet 5. Ordinary bug fixing is unaffected, and verified security teams can apply for broader access.
  • Not a drop-in API swap: Anthropic lists five breaking changes from Sonnet 5, including errors on forced tool use and on non-default temperature settings. Read the migration guide before switching production code.
03

Benchmark Snapshot

Terminal-Bench 4.0 — 64% independent (70.6% vendor)

Multi-step work in a command-line terminal. Artificial Analysis measured 64%, against about 60% for Opus 5.5 and GPT-6 Astra (xhigh). Anthropic's own best run is 70.6%, versus 66.4% for Opus 5.5 at xhigh and 10.3% for Sonnet 5.

FrontierCode v1.1 — 52.1% at xhigh, 46.2% at max

Whether code changes are ready to merge. Anthropic's table: Opus 5.5 54.4%, GPT-6 Sol 49.3%, Sonnet 5 42.4%. Sonnet 5.5 scores lower at max than at xhigh.

CursorBench 4.0 — 55.5%

Ambiguous, multi-file tasks taken from real Cursor sessions. Anthropic reports Opus 5.5 at 57.8% and Sonnet 5 at 34.1%.

Vals Index — 69.22% (2nd of 66)

Independent index from Vals AI, 0.47 points behind Opus 5.5 and ahead of Fable 5.1, at $20.80 per test versus $32.77 for Opus 5.5. First on Vibe Code Bench and Code Migration.

Output tokens per task — ~193k at max effort

Artificial Analysis measurement across its Intelligence Index, the highest it has recorded: about 60% above Opus 5.5 (max) and about seven times GPT-6 Astra (max).

Price — $2 / $10 per 1M tokens, $0.20 cache read

Unchanged from Sonnet 5 and half of Opus 5.5's $4/$20. Cache reads cost the same $0.20 as Opus, so cache-heavy coding sessions save less than the headline price suggests. Batch requests are 50% off.

04

The Verdict

Claude Sonnet 5.5 is the model to leave running for everyday engineering: scoped bugs, features, migrations, and front-end work, at medium or high effort, where it is fast and cheap. It ranks just behind GPT-6 Astra and Opus 5.5 because its best independent terminal score comes at max effort, the setting where it burns the most tokens and where its merge-quality score falls. Keep it below max, and hand the open-ended architecture questions to Opus 5.5.

05

Frequently Asked Questions