Ranked #6 Coding — AI That Writes Production Code
xAI

Grok 4.7

Grok 4.7 is the cheapest serious coding model near the top of the leaderboard. At $2 in and $6 out per million tokens, it scores 71.0% on DeepSWE and 46.3% on CursorBench 4.0, and it runs natively in Cursor and Grok Build. It's a strong everyday pair programmer, though it's not yet the agent you'd leave alone in a terminal for hours.

Updated September 27, 2026 DeepSWE 71.0%CursorBench 46.3%Cursor & Grok Build
9.7out of 10
Official Website
Best for

Grok 4.7 is the cheapest serious coding model near the top of the leaderboard. At $2 in and $6 out per million tokens, it scores 71.0% on DeepSWE and 46.3% on CursorBench 4.0, and it runs natively in Cursor and Grok Build. It's a strong everyday pair programmer, though it's not yet the agent you'd leave alone in a terminal for hours.

Why It Wins

In xAI's launch table: CursorBench 4.0 rises from 40.4% to 46.3%, DeepSWE v1.1 from 65.2% to 71.0% (high effort), and Terminal-Bench 4.0 nearly doubles from 20.3% to 37.6%. Price is unchanged at $2/$6, with a fast variant at twice the speed for twice the price. Available in Cursor, Grok Build (free to try), the Grok API, and third-party coding tools.

Watch out

On long, autonomous terminal work it still trails the leaders badly: 37.6% on Terminal-Bench 4.0 versus 57.9% for Claude Fable 5.1 and 59.6% for GPT-6 Astra and Claude Opus 5.5 in independent testing. On CursorBench it sits behind Fable 5.1 (51.8%). All Grok figures are xAI's own.

01

What It Actually Is

Every developer has worked next to two kinds of colleague.

One is brilliant and thorough, and gives you a fifteen-minute lecture every time you ask a question. The other just leans over, types forty lines, runs the tests, and says “try that.”

In OpenAI and Anthropic’s line-ups, the careful, expensive colleague has the most titles. Grok 4.7 is the quick one who is cheap enough to ask a hundred times a day.

Released on September 21, 2026, Grok 4.7 sits at #5 on our Coding leaderboard. Its role is to be the best-value pair programmer in the top tier.

The improvements that matter to coders

Grok 4.6 was an interesting challenger. Grok 4.7 is a clear step up on every coding test xAI published:

  • DeepSWE v1.1, complex engineering tasks in real codebases: from 65.2% to 71.0% (at high effort). In xAI’s table that’s slightly ahead of Claude Fable 5.1’s 70.0%.
  • CursorBench 4.0, longer coding tasks inside the Cursor editor: from 40.4% to 46.3%.
  • Terminal-Bench 4.0, multi-hour work in a command line: from 20.3% to 37.6%, nearly double.

xAI says the gains come from a larger base model, much longer training on tasks that take hours to finish, and specific work on checking its own answers and handling long context. In practice, that means fewer half-finished changes and fewer confident mistakes.

The economics of coding in the editor

When you code in Cursor or a similar editor, the AI isn’t called once a day. It’s called constantly. It reads your files, suggests edits, explains an error, rewrites a function, and runs again.

At $2 per million input tokens and $6 per million output, Grok 4.7 has the lowest output price among top-tier coding models. For comparison: GPT-6 Sol is $10, Claude Opus 5.5 is $20, and GPT-6 Astra is $50. If you work in the editor all day, that difference is real money.

There’s also a fast variant that produces output twice as fast at twice the price. It’s useful when you’re waiting on every answer.

Where it still falls short

Long unsupervised runs. Terminal-Bench 4.0 is the test that best predicts whether you can leave an agent alone in a terminal for hours. Grok 4.7’s 37.6% is a big improvement, but GPT-6 Astra and Claude Opus 5.5 both score 59.6% in independent testing, and Fable 5.1 scores 57.9% in xAI’s own table. For overnight autonomous work, those models are much less likely to get lost.

The best editor results. On CursorBench 4.0, Grok trails Claude Fable 5.1 (51.8%) and Anthropic’s reported 57.8% for Claude Opus 5.5.

Newest rivals. xAI compares Grok 4.7 with GPT-5.6 Sol and Fable 5.1. GPT-6 Sol launched the next day at a similar price, and we haven’t seen an independent head-to-head yet. All Grok numbers here are xAI’s own.

The verdict

For long, unsupervised agent runs, choose GPT-6 Astra. For deep architectural work across a large codebase, choose Claude Opus 5.5.

But for the hours you spend each day in your editor, fixing bugs, writing tests, and refactoring components, Grok 4.7 is a strong, fast, and very cheap partner. Test it against GPT-6 Sol on your own code; one of the two will likely become your daily driver.

02

Strengths and honest limitations

Key Strengths

  • The lowest price near the top: $2 per million input tokens and $6 per million output. That’s cheaper on output than GPT-6 Sol ($10), Claude Opus 5.5 ($20), and GPT-6 Astra ($50). In an editor that calls the model constantly, it adds up.
  • Solid real-world bug fixing: On DeepSWE v1.1, which tests complex engineering tasks in real codebases, Grok 4.7 scored 71.0% at high effort, up from 65.2% on Grok 4.6. By xAI’s figures that’s slightly ahead of Claude Fable 5.1 (70.0%).
  • Better at editor work: CursorBench 4.0 measures longer-running coding tasks inside the Cursor editor. Grok 4.7 improved from 40.4% to 46.3%, and xAI calls it ‘at the frontier in price-performance’ on that test.
  • Stays on task longer: xAI trained it on tasks that take many hours and improved how it checks its own work and handles long context. Those are exactly the skills that stop an agent from wandering off halfway through a change.
  • Available where developers work: Built into Cursor and Grok Build on day one, and available through the Grok API, other coding tools, and model routers. A fast variant doubles output speed for interactive use.

Honest Limitations

  • Not ready for long unsupervised terminal runs: Terminal-Bench 4.0, which tests multi-hour command-line work, rose to 37.6%. That’s a big jump, but still far behind the roughly 58–60% of Fable 5.1, GPT-6 Astra, and Claude Opus 5.5.
  • Behind the leaders on editor tasks: Its 46.3% on CursorBench 4.0 is below Fable 5.1 (51.8%) and Anthropic’s reported 57.8% for Claude Opus 5.5.
  • Compared against last month’s rivals: xAI’s table sets Grok 4.7 against GPT-5.6 Sol and Fable 5.1, not GPT-6 Sol, which launched the next day at a similar price.
  • Vendor numbers only: Every Grok figure here comes from xAI. We’ll update when independent results arrive.
03

Benchmark Snapshot

DeepSWE v1.1 — 71.0% (high effort)

Complex engineering tasks in real codebases. Up from 65.2% on Grok 4.6. xAI's table shows Claude Fable 5.1 at 70.0% and GPT-5.6 Sol (max) at 72.7%.

CursorBench 4.0 — 46.3%

Longer-running coding tasks in the Cursor editor. Up from 40.4%; Claude Fable 5.1 scores 51.8% and GPT-5.6 Sol 41.7% in the same table.

Terminal-Bench 4.0 — 37.6%

Multi-hour command-line work. Nearly double Grok 4.6's 20.3%, but well behind Claude Fable 5.1's 57.9% in xAI's table. (Anthropic's own run puts Fable 5.1 at 55.8%; independent testing puts GPT-6 Astra and Claude Opus 5.5 at 59.6%.)

Price — $2 / $6 per 1M tokens

Unchanged from Grok 4.6. The fast variant runs at twice the output speed for twice the price.

04

The Verdict

Grok 4.7 is our #5 coding model and the best choice if cost per keystroke is what you care about most. It’s a genuinely capable pair programmer for everyday work in the editor, like fixing bugs, writing tests, and refactoring, at the lowest output price in the top tier. For long unsupervised terminal sessions, it still falls well short of GPT-6 Astra and Claude Opus 5.5. For the closest price rival, test it against GPT-6 Sol on your own code before you commit.

05

Frequently Asked Questions