Picture two carpenters. One is the master you call when a house needs a new staircase and nobody has drawn the plans. The other is the steady tradesperson who can hang forty doors in a week, every one of them square.
Most of a software team’s week is doors, not staircases: a failing test, a form that needs a new field, an API that changed its name. Claude Sonnet 5.5 is built for the doors.
Anthropic released it on September 28, 2026, as the second model in the Claude 5.5 family, a week after Opus 5.5. It keeps Sonnet 5’s price of $2 per million input tokens and $10 per million output, which is half of Opus 5.5. Anthropic says it runs more than 30% faster than Sonnet 5 and costs up to 30% less per task because it finishes in fewer steps.
The independent numbers
For coding agents, the test we trust most is Terminal-Bench 4.0: long, multi-step jobs in a real command-line environment, where the model has to install things, run them, read the errors, and try again.
Artificial Analysis ran it independently in September 2026. Sonnet 5.5 scored 64%, slightly ahead of Opus 5.5 and GPT-6 Astra at about 60% each. Anthropic’s own run is higher, at 70.6%, but vendor harnesses tend to flatter their own models, so the independent 64% is the number to remember. Either way, it is an enormous jump from Sonnet 5, which scored 10.3% in Anthropic’s table.
A second independent source agrees on the big picture. Vals AI places Sonnet 5.5 second of 66 models on its index at 69.22%, less than half a point behind Opus 5.5, while costing about two-thirds as much per test. It takes first place on two coding leaderboards: Vibe Code Bench (building small apps from a description) at 92.39% and Code Migration at 69.83%.
What Anthropic’s own tests add
Anthropic’s launch table fills in the details:
- FrontierCode v1.1, which asks whether a code change could be merged without a human touching it: 52.1% at xhigh effort. Opus 5.5 scores 54.4% and GPT-6 Sol 49.3%.
- CursorBench 4.0, built from real sessions in the Cursor editor: 55.5%, about two points behind Opus 5.5 and far ahead of Sonnet 5’s 34.1%.
The more useful part is the cost charts, which plot every model’s score against its cost per task at each effort level. Two findings stand out. At High effort, the API default, Sonnet 5.5 matches GPT-6 Sol’s best FrontierCode score for about a fifth of the cost per task. And at Medium effort, the default in Claude Code, it beats Sonnet 5’s best Terminal-Bench score for less than a tenth of the cost.
Early testers describe the same thing in plainer words. Lovable saw roughly half as many shell commands per task. Slack measured about 14% fewer output tokens than Sonnet 5. Unity, which only counts a task as done when the change actually works at runtime, says the majority of Sonnet 5.5’s work passed that check.
The effort dial, and why the top setting backfires
Like Opus 5.5, Sonnet 5.5 has an effort dial with five settings: low, medium, high, xhigh, and max. Higher settings let the model think longer and check its work more. You would expect max to be best. For Sonnet 5.5, it often isn’t.
Artificial Analysis measured the whole range on its Intelligence Index:
| Effort | Intelligence Index | Cost per task |
|---|---|---|
| Low | 36 | $0.41 |
| Medium | 41 | $0.59 |
| High | 47 | $1.08 |
| Xhigh | 52 | $2.74 |
| Max | 56 | $7.60 |
Going from xhigh to max buys four more points for nearly three times the cost. At max, Sonnet 5.5 wrote about 193,000 output tokens per task. That is the most Artificial Analysis has ever measured: about 60% more than Opus 5.5 at max, and about seven times GPT-6 Astra.
Worse, more effort does not always mean better code. On FrontierCode, Sonnet 5.5 scores 52.1% at xhigh but 46.2% at max. Anthropic explains why in a footnote: at max, the model more often launched Claude Code’s code-review routine, which splits the review across many helper agents. In some cases that timed out or made extra edits beyond the task, and FrontierCode penalizes changes nobody asked for.
The lesson is simple: treat xhigh as the ceiling. If a task still fails at xhigh, the better next step is usually Opus 5.5, not max.
The honest catch
Opus is still the architect. Anthropic says it plainly: on several tests Sonnet 5.5 at max comes close to Opus 5.5, but in Anthropic’s own testing and that of outside testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment. Opus also keeps the lead on FrontierCode and CursorBench.
The Claude Code bill won’t halve. Coding agents re-read your project constantly, and those re-reads are billed as cache reads, which cost $0.20 per million tokens on both Sonnet 5.5 and Opus 5.5. You save on the new input and output, but a long cache-heavy session saves less than the headline “half price” suggests.
Security work is rerouted. Sonnet 5.5 is the first Sonnet to ship with the cyber safeguards Anthropic uses on its top models. Routine bug fixing is unaffected, but higher-risk security tasks visibly fall back to Sonnet 5 unless your team is approved through Anthropic’s Cyber Verification Program.
Migration takes a little work. Anthropic lists five breaking changes from Sonnet 5. Forced tool use now returns an error, non-default temperature settings are rejected, and text between tool calls comes back in a different format. None of this is hard to fix, but it is not a one-line model swap.
Where it fits
| If the job is… | Reach for | Why |
|---|---|---|
| A scoped bug, feature, or migration | Claude Sonnet 5.5 | Fast and cheap at medium or high effort |
| Open-ended architecture or a problem nobody has scoped | Claude Opus 5.5 | Clearly stronger judgment, per Anthropic and testers |
| Long unattended terminal work at the lowest cost per task | GPT-6 Astra | Uses a fraction of the tokens at peak |
| Front-end polish, slides, small visual builds | Claude Sonnet 5.5 | A sharp eye for design and fast iteration |
The verdict
Claude Sonnet 5.5 is the model most developers should leave switched on. It handles the doors, all forty of them, quickly and at half the flagship’s token price. Keep it at medium or high effort, use xhigh for the hard ones, and when the job turns out to be a staircase, call Opus.