Ranked #7 Everyday Ecosystem — The Leading AI Assistants
xAI

Grok 4.7

Grok 4.7 keeps the same low price as Grok 4.6 — $2 in, $6 out per million tokens — but is a noticeably stronger, more careful model. xAI trained it on longer, harder tasks and taught it to check its own work. It's the best-value way to get near-frontier help with office work, legal drafts, and engineering maths.

Updated September 27, 2026 Value Frontier$2/$6 PricingSelf-Verification
9.6out of 10
Official Website
Best for

Grok 4.7 keeps the same low price as Grok 4.6 — $2 in, $6 out per million tokens — but is a noticeably stronger, more careful model. xAI trained it on longer, harder tasks and taught it to check its own work. It's the best-value way to get near-frontier help with office work, legal drafts, and engineering maths.

Why It Wins

Same $2/$6 price as Grok 4.6, with a fast variant at twice the speed for twice the price. In xAI's launch table, Grok 4.7 has the top score on the Harvey Legal Agent Benchmark (19.6%, versus 6.7% for Claude Fable 5.1) and EEBench electrical engineering (64.0%). It lands close to Fable 5.1 on AA Briefcase multi-hour office work (1,657 vs 1,678). xAI calls it its strongest model yet at refusing harmful requests and resisting jailbreaks.

Watch out

Its launch table compares against GPT-5.6 Sol and Claude Fable 5.1, not the newer Claude Opus 5.5 or GPT-6 Sol, so its position against today's best is not yet settled. It trails clearly on long terminal work (37.6% on Terminal-Bench 4.0 versus 57.9% for Fable 5.1) and on clinical reasoning. All figures are xAI's own.

01

What It Actually Is

Imagine a budget airline that keeps its fares the same but quietly swaps in a bigger, better plane.

That’s roughly what xAI did with Grok 4.7, released on September 21, 2026. The price didn’t change: $2 per million input tokens and $6 per million output, exactly the same as Grok 4.6. What changed is the model underneath.

What’s new under the hood

xAI describes three changes, in plain terms.

  1. A larger base model. xAI says 4.7 is built on a bigger foundation than 4.6. It hasn’t published a parameter count in its launch post, so treat any specific figure you see elsewhere as unconfirmed.
  2. Longer, harder training. The model was trained with reinforcement learning on tasks that take many hours to finish, which is the kind of work where earlier models lose the thread halfway through.
  3. Checking its own work. xAI trained it specifically to verify its answers and to handle longer documents more reliably.

The result is a model that’s better at sticking with a long task and catching its own mistakes before you do.

Where it surprised us

The most interesting numbers in xAI’s launch table aren’t the coding ones. They’re in areas where you wouldn’t expect a budget model to lead.

  • Legal work. On the Harvey Legal Agent Benchmark, a test of how well an AI agent handles real legal tasks, Grok 4.7 scored 19.6%. That’s ahead of Claude Fable 5.1 (6.7%) and GPT-5.6 Sol (2.5%). It’s a brutally hard test and everyone scores low, but Grok came out on top of that comparison.
  • Engineering maths. On EEBench, which tests electrical engineering problems, it scored 64.0%, ahead of Fable 5.1 at 56.4%.
  • Office work. On AA Briefcase, a test of multi-hour office tasks like building documents and presentations, it scored 1,657, just behind Fable 5.1’s 1,678 and far above Grok 4.6’s 1,546.

Put simply: on real office and professional tasks, Grok 4.7 is now in the same conversation as models that cost several times more.

Safer than its reputation

Grok has a reputation as the chatbot with fewer guardrails. For 4.7, that reputation is out of date. xAI built what it calls an entirely new safety system and says this is its strongest model yet at refusing dangerous requests and resisting jailbreaks. It also claims the model rarely blocks legitimate work, such as defensive security research.

Grok still tends to answer directly rather than wrapping replies in disclaimers. That’s a style choice, though, not an absence of safety.

The honest catch

It hasn’t met the newest rivals yet. xAI’s table compares Grok 4.7 with GPT-5.6 Sol and Claude Fable 5.1. A day after it launched, Anthropic released Claude Opus 5.5 and OpenAI released GPT-6 Sol, and both are strong, well-priced competitors. Until independent tests put them side by side, Grok’s exact position is uncertain.

Long terminal work is still a weakness. On Terminal-Bench 4.0, which tests multi-hour work in a command line, Grok 4.7 scored 37.6%. That’s nearly double Grok 4.6, but well behind Fable 5.1’s 57.9%.

Medicine is not its strength. On HealthBench Professional, a clinical reasoning test, it scored 56.7%, below both GPT-5.6 Sol and Fable 5.1.

These are all xAI’s numbers. The legal and engineering leads are exciting, but we’d like to see them confirmed independently.

Who should use it

If you use AI all day for documents, analysis, and technical questions, and the monthly bill matters, Grok 4.7 is worth a serious trial. It gives you most of the quality of the top models at a fraction of the cost.

If you need the most careful judgment available, go with Claude Opus 5.5. If you want an agent that operates software on your computer, go with GPT-6 Astra. For everything in between, Grok 4.7 is the best value near the frontier.

02

Strengths and honest limitations

Key Strengths

  • Frontier-class work at a budget price: $2 per million input tokens and $6 per million output. That’s under a third of Claude Opus 5.5’s output price and about an eighth of GPT-6 Astra’s. For teams that use AI all day, the difference shows up on the monthly bill.
  • Surprisingly strong at legal work: On the Harvey Legal Agent Benchmark, Grok 4.7 scored 19.6% in xAI’s table, ahead of Claude Fable 5.1 (6.7%) and GPT-5.6 Sol (2.5%). It’s a hard test where every model still scores low, but Grok led the comparison.
  • Good with engineering maths: On EEBench, a test of electrical engineering problems, it scored 64.0%, ahead of Fable 5.1 (56.4%) and well up from Grok 4.6 (53.0%).
  • Checks its own work: xAI trained Grok 4.7 on longer, harder tasks and specifically improved how it verifies its answers and handles long documents. On AA Briefcase, a test of multi-hour office work, it jumped from 1,546 to 1,657, close to Fable 5.1’s 1,678.
  • Safer without being preachy: xAI rebuilt the safety system. It calls Grok 4.7 its strongest model yet at refusing dangerous requests and resisting jailbreaks, while rarely blocking legitimate work such as defensive security research.

Honest Limitations

  • Not compared against the newest rivals: xAI’s table benchmarks Grok 4.7 against GPT-5.6 Sol and Claude Fable 5.1. Claude Opus 5.5 and GPT-6 Sol arrived a day later, so independent head-to-head results are still needed.
  • Weaker at long terminal sessions: On Terminal-Bench 4.0, which tests multi-hour work in a command line, it scored 37.6%, a big jump from Grok 4.6’s 20.3% but far behind Fable 5.1 at 57.9%.
  • Behind on medicine: On HealthBench Professional, a test of clinical reasoning, it scored 56.7%, below GPT-5.6 Sol (60.5%) and Fable 5.1 (62.1%).
  • Vendor numbers only: Every figure here comes from xAI’s launch post. Treat the legal and engineering leads as promising until independent tests confirm them.
03

Benchmark Snapshot

Harvey Legal Agent Benchmark — 19.6%

Agentic legal work. xAI's table: Grok 4.7 19.6%, Grok 4.6 15.8%, Claude Fable 5.1 6.7%, GPT-5.6 Sol 2.5%. A very hard test; low absolute scores are normal.

EEBench — 64.0%

Electrical engineering problem solving. Grok 4.6 scored 53.0%, Claude Fable 5.1 56.4%, GPT-5.6 Sol 39.4% in the same table.

AA Briefcase v1.1 — 1,657

Multi-hour office tasks such as documents and presentations. Up from 1,546 on Grok 4.6 and just behind Claude Fable 5.1 at 1,678.

Terminal-Bench 4.0 — 37.6%

Multi-hour terminal work. A large improvement over Grok 4.6 (20.3%), level with GPT-5.6 Sol (37.3%), but well behind Claude Fable 5.1 (57.9%) in xAI's table. (Anthropic's own run puts Fable 5.1 at 55.8%; independent testing puts GPT-6 Astra and Claude Opus 5.5 at 59.6%.)

Price — $2 / $6 per 1M tokens

Unchanged from Grok 4.6. A fast variant offers twice the output speed at twice the price.

04

The Verdict

Grok 4.7 holds our #5 Everyday Ecosystem spot as the best-value model near the frontier. It isn’t the smartest assistant you can buy. Claude Opus 5.5 and GPT-6 Astra still lead on careful judgment and operating software. But Grok 4.7 gets surprisingly close on office work, leads its launch table on legal and engineering tests, and costs a fraction as much. If you use AI all day and watch the bill, it deserves a serious trial.

05

Frequently Asked Questions