Imagine a budget airline that keeps its fares the same but quietly swaps in a bigger, better plane.
That’s roughly what xAI did with Grok 4.7, released on September 21, 2026. The price didn’t change: $2 per million input tokens and $6 per million output, exactly the same as Grok 4.6. What changed is the model underneath.
What’s new under the hood
xAI describes three changes, in plain terms.
- A larger base model. xAI says 4.7 is built on a bigger foundation than 4.6. It hasn’t published a parameter count in its launch post, so treat any specific figure you see elsewhere as unconfirmed.
- Longer, harder training. The model was trained with reinforcement learning on tasks that take many hours to finish, which is the kind of work where earlier models lose the thread halfway through.
- Checking its own work. xAI trained it specifically to verify its answers and to handle longer documents more reliably.
The result is a model that’s better at sticking with a long task and catching its own mistakes before you do.
Where it surprised us
The most interesting numbers in xAI’s launch table aren’t the coding ones. They’re in areas where you wouldn’t expect a budget model to lead.
- Legal work. On the Harvey Legal Agent Benchmark, a test of how well an AI agent handles real legal tasks, Grok 4.7 scored 19.6%. That’s ahead of Claude Fable 5.1 (6.7%) and GPT-5.6 Sol (2.5%). It’s a brutally hard test and everyone scores low, but Grok came out on top of that comparison.
- Engineering maths. On EEBench, which tests electrical engineering problems, it scored 64.0%, ahead of Fable 5.1 at 56.4%.
- Office work. On AA Briefcase, a test of multi-hour office tasks like building documents and presentations, it scored 1,657, just behind Fable 5.1’s 1,678 and far above Grok 4.6’s 1,546.
Put simply: on real office and professional tasks, Grok 4.7 is now in the same conversation as models that cost several times more.
Safer than its reputation
Grok has a reputation as the chatbot with fewer guardrails. For 4.7, that reputation is out of date. xAI built what it calls an entirely new safety system and says this is its strongest model yet at refusing dangerous requests and resisting jailbreaks. It also claims the model rarely blocks legitimate work, such as defensive security research.
Grok still tends to answer directly rather than wrapping replies in disclaimers. That’s a style choice, though, not an absence of safety.
The honest catch
It hasn’t met the newest rivals yet. xAI’s table compares Grok 4.7 with GPT-5.6 Sol and Claude Fable 5.1. A day after it launched, Anthropic released Claude Opus 5.5 and OpenAI released GPT-6 Sol, and both are strong, well-priced competitors. Until independent tests put them side by side, Grok’s exact position is uncertain.
Long terminal work is still a weakness. On Terminal-Bench 4.0, which tests multi-hour work in a command line, Grok 4.7 scored 37.6%. That’s nearly double Grok 4.6, but well behind Fable 5.1’s 57.9%.
Medicine is not its strength. On HealthBench Professional, a clinical reasoning test, it scored 56.7%, below both GPT-5.6 Sol and Fable 5.1.
These are all xAI’s numbers. The legal and engineering leads are exciting, but we’d like to see them confirmed independently.
Who should use it
If you use AI all day for documents, analysis, and technical questions, and the monthly bill matters, Grok 4.7 is worth a serious trial. It gives you most of the quality of the top models at a fraction of the cost.
If you need the most careful judgment available, go with Claude Opus 5.5. If you want an agent that operates software on your computer, go with GPT-6 Astra. For everything in between, Grok 4.7 is the best value near the frontier.