Ranked guide

Coding — AI That Writes Production Code

These are coding agents, not autocomplete toys. GPT-6 Astra leads because it finishes terminal work in the fewest tokens, while Claude Opus 5.5 is level with it on independent tests and strongest on big, messy codebases. Claude Sonnet 5.5 is the fast, half-price option for well-scoped bugs and features. GPT-6.1 Sol and Grok 4.7 bring near-frontier coding at a fraction of the price, and GLM-5.3, GLM-5.3-Flash and Qwen3.8-Max round out the shortlist with strong value options.

Decision first

Our ranking

Start with the winner, then compare the trade-offs that might change the answer for you.

#1 Coding

GPT-6 Astra

OpenAI

GPT-6 Astra takes the coding crown the same honest way GPT-5.6 held it: by finishing agent-shaped work — terminal sessions, failing pipelines, incident forensics, whole-desktop setups — now backed by Arena.ai's new #1 ranking on Code Arena (WebDev 1,797 pts) while burning roughly a third of Sol's tokens to do it.

Why It Wins

New #1 on Arena.ai's Code Arena: WebDev with 1,797 pts (+35 lead over Claude Fable 5.1 Max at 1,762, +109 over Opus 5 Max at 1,688, and +180 over Sol at #13), leading Data & Analytics, Consumer Product, and Content Creation Tools. Combined with Terminal-Bench 4.0 at 57.7–57.9% (Fable 5.1: 55.8%), table-best DeepSWE v1.1 at 74.1%, Terminal-Bench Science at 64.6%, SRE-Bench at 88% pass@1 / 99.2% pass@4, and 100% on ExploitBench with previously unknown 0-days found during evaluation. While Artificial Analysis' own agent harness still narrowly scores Fable 5.1 ahead (70 vs 67), Astra reshapes the Pareto frontier at $40/Mtoken blended and burns roughly a third of Sol's tokens.

The Catch

Even with Code Arena's #1 crown in WebDev and consumer tooling, Fable 5(.1) still holds the SWE-Bench Pro repo-issue record and leads on Artificial Analysis' own-harness agent index (70 vs 67). The API sticker is 2.5x Sol at $10/$50, cache reads cost four times Fable 5.1's $0.25, terse answers can skip report-polish steps, and stronger cyber safeguards add friction to exploit-adjacent work. Launch tables moved after publish — treat any single number as ±1–2 points.

9.9 Editorial score
Read review
Best for

GPT-6 Astra takes the coding crown the same honest way GPT-5.6 held it: by finishing agent-shaped work — terminal sessions, failing pipelines, incident forensics, whole-desktop setups — now backed by Arena.ai's new #1 ranking on Code Arena (WebDev 1,797 pts) while burning roughly a third of Sol's tokens to do it.

Why It Wins

New #1 on Arena.ai's Code Arena: WebDev with 1,797 pts (+35 lead over Claude Fable 5.1 Max at 1,762, +109 over Opus 5 Max at 1,688, and +180 over Sol at #13), leading Data & Analytics, Consumer Product, and Content Creation Tools. Combined with Terminal-Bench 4.0 at 57.7–57.9% (Fable 5.1: 55.8%), table-best DeepSWE v1.1 at 74.1%, Terminal-Bench Science at 64.6%, SRE-Bench at 88% pass@1 / 99.2% pass@4, and 100% on ExploitBench with previously unknown 0-days found during evaluation. While Artificial Analysis' own agent harness still narrowly scores Fable 5.1 ahead (70 vs 67), Astra reshapes the Pareto frontier at $40/Mtoken blended and burns roughly a third of Sol's tokens.

Watch out

Even with Code Arena's #1 crown in WebDev and consumer tooling, Fable 5(.1) still holds the SWE-Bench Pro repo-issue record and leads on Artificial Analysis' own-harness agent index (70 vs 67). The API sticker is 2.5x Sol at $10/$50, cache reads cost four times Fable 5.1's $0.25, terse answers can skip report-polish steps, and stronger cyber safeguards add friction to exploit-adjacent work. Launch tables moved after publish — treat any single number as ±1–2 points.

#2

Claude Opus 5.5

Anthropic

Claude Opus 5.5 is the coding model to call when the job is big and tangled: a migration across hundreds of thousands of lines, a bug that crosses five services, a legacy system nobody fully understands anymore. On Anthropic's own benchmarks it leads GPT-6 Astra; on independent testing the two are tied. It holds our #2 spot only because Astra finishes the same terminal work in far fewer tokens.

9.9 Editorial score
Read review
#3

Claude Sonnet 5.5

Anthropic

Claude Sonnet 5.5 is the coding model you can leave switched on all day: fast, half the token price of Opus 5.5, and good enough that most bug fixes, features, and refactors never need the flagship. Just don't turn the effort dial all the way up — at its highest setting it writes more than any model Artificial Analysis has measured, and some results get worse.

9.8 Editorial score
Read review
#4

GPT-6.1 Sol

OpenAI

GPT-6.1 Sol is the coding agent you can leave running all day without watching the bill. On independent tests it lands within a point or two of GPT-6 Astra, OpenAI's flagship, yet a typical task costs less than a quarter as much. The token price hasn't changed from GPT-6 Sol ($2 in, $10 out per million), but re-reading cached code now costs just $0.10 per million tokens, and big repositories are exactly where that saving shows.

9.8 Editorial score
Read review
#5

Claude Fable 5.1

Anthropic

The Mythos-class reasoning engine refined for agentic coding. Same weights as the restricted Mythos 5.1, but optimized with a 75% cache discount that transforms long-horizon engineering economics. Best deployed in Claude Code, where its AA Coding Agent score of 70 leads the field.

9.8 Editorial score
Read review
#6

Grok 4.7

xAI

Grok 4.7 is the cheapest serious coding model near the top of the leaderboard. At $2 in and $6 out per million tokens, it scores 71.0% on DeepSWE and 46.3% on CursorBench 4.0, and it runs natively in Cursor and Grok Build. It's a strong everyday pair programmer, though it's not yet the agent you'd leave alone in a terminal for hours.

9.7 Editorial score
Read review
#7

GLM-5.3

Z.ai (Zhipu AI)

A heavyweight open-weight coding agent trained for the part that comes after the first answer: planning, editing, testing, recovering, and staying oriented through long repository jobs.

9.2 Editorial score
Read review
#8

Qwen3.8-Max

Alibaba / Qwen Team

A high-value visual and long-context coding challenger, now ranked #6 after independent testing placed flagship GLM-5.3 ahead. Its downloadable Max-tier checkpoint remains remarkable, but 2.4T parameters make self-hosting a datacenter project.

9.2 Editorial score
Read review
#9

GLM-5.3-Flash

Z.ai (Zhipu AI)

Frontier-adjacent coding and agent work at unusually low prices, with native vision and a 1M context. “Flash” means efficient and cheap here—not small, and not especially fast.

9.2 Editorial score
Read review
Questions, answered

Frequently Asked Questions