Imagine an office with two senior people. One is the partner you call into the meeting when the deal is unusual and the stakes are high. The other is the colleague who gets through the week’s pile, forty emails, three spreadsheets, a slide deck for Thursday, quickly and neatly.
Most of us need the second person far more often than the first. Claude Sonnet 5.5 is that colleague.
Anthropic released it on September 28, 2026, a week after Claude Opus 5.5, as the second model in the Claude 5.5 family. It is meant as the everyday complement to Opus: in Anthropic’s words, Opus is built for complex work requiring careful judgment, while Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets.
How close is it to Opus?
Closer than you might expect, at least on the independent tests.
Artificial Analysis runs two tests built around finished professional work rather than quiz questions. In GDPval-AA v2.1, models produce real deliverables, such as a financial model, a legal memo, or a briefing, and the results are compared head-to-head like chess games. Sonnet 5.5 scored 1844 Elo. Opus 5.5 scored 1846, a gap too small to matter. On the newer AA-Briefcase, which tests longer multi-step knowledge work, it scored 1811 against Opus’s 1822.
The broadest independent measure is the Artificial Analysis Intelligence Index, a composite of hard evaluations. At max effort, Sonnet 5.5 scored 56, second only to Opus 5.5 at 58.
That is a huge step for a Sonnet. Its predecessor, Sonnet 5, scored 1449 on GDPval-AA. In practical terms, office work that used to need the flagship can now go to the cheaper, faster model.
Where Opus still pulls ahead
There are two places where the gap is real.
Facts. Artificial Analysis’s AA-Omniscience test asks thousands of factual questions. Sonnet 5.5 answered 54% correctly; Opus 5.5 answered 66%. A smaller model simply holds less knowledge. There is a silver lining: when Sonnet 5.5 didn’t know an answer, it made one up 47% of the time, compared with 59% for Opus 5.5. It knows less, but it is more willing to say so. Still, for an obscure date, a name, or a statistic, ask it to search the web or cite a source.
Judgment. Anthropic is unusually direct about this. On several tests, Sonnet 5.5 at max effort performs comparably to Opus 5.5. But in Anthropic’s own testing and that of outside testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment. If a question has no obvious frame, such as a strategy decision or a contract with unusual terms, that is still Opus territory.
What it’s good at
Anthropic highlights documents, summaries, spreadsheets, and slides, and says the model has a sharp eye for design: it can follow a slide template closely enough that little editing is needed. Its chart reading improved dramatically. On Anthropic’s Chartography test, which asks models to read values from charts without tools, it scored 61.6%, up from 15.6% for Sonnet 5.
It is also quick. Anthropic says output arrives more than 30% faster than Sonnet 5, and early testers noticed: Zendesk says support tickets were processed 20% faster, and Box measured 2.4 times the speed of the previous version.
It can operate a computer, too. On OSWorld 2.1, where a model clicks and types its way through real desktop apps, Anthropic reports 80.1% under partial-credit scoring, close to Opus 5.5’s 81.8%.
Leave the dial in the middle
Sonnet 5.5 has five effort settings: low, medium, high, xhigh, and max. The Claude apps use medium by default. Here is what Artificial Analysis measured at each:
| Effort | Intelligence Index | Cost per task |
|---|---|---|
| Low | 36 | $0.41 |
| Medium | 41 | $0.59 |
| High | 47 | $1.08 |
| Xhigh | 52 | $2.74 |
| Max | 56 | $7.60 |
The last step is the expensive one. At max, Sonnet 5.5 wrote about 193,000 output tokens per task, the most Artificial Analysis has ever measured, and about 60% more than Opus 5.5 at max. For most people in the Claude app this simply means waiting longer. For developers paying per token, it means that a maxed-out Sonnet is no longer the cheap option. If a task needs that much thinking, it usually deserves Opus.
The honest catch
No pictures. Claude reads screenshots, charts, and scanned PDFs well, but it doesn’t generate images, audio, or video. If you want one app that also draws and talks, ChatGPT and Gemini are broader.
Price per token is not price per task. Tokens cost the same as Sonnet 5, $2 per million in and $10 per million out. Anthropic says typical tasks cost up to 30% less because the model finishes in fewer steps. That holds at normal settings; at max, as the table shows, the bill climbs fast.
Some requests are guarded. Sonnet 5.5 is the first Sonnet to ship with the stricter cybersecurity safeguards Anthropic uses on its top models. Higher-risk security requests visibly fall back to Sonnet 5. For almost everyone, this never comes up.
Who should use it
| If your day looks like… | Reach for | Why |
|---|---|---|
| Emails, summaries, spreadsheets, slide decks | Claude Sonnet 5.5 | Near-Opus office work, fast, on the free plan |
| High-stakes judgment, obscure facts, messy research | Claude Opus 5.5 | More knowledge and clearly stronger judgment |
| “Operate this software on my computer and finish the job” | GPT-6 Astra | Stronger, more economical computer-use agent |
| One app for chat, images, and voice | ChatGPT or Gemini | Broader creative toolbox |
The everyday verdict
Claude Sonnet 5.5 won’t replace Opus for the hardest questions, and Anthropic doesn’t pretend it will. But the hardest questions are a small part of most weeks. For everything else, it is fast, capable, available on Claude’s free plan, and priced so you don’t have to think about it. Leave it on the default setting, check the obscure facts, and call in Opus when the stakes go up.