Picture a building site with two crane operators.
One is the veteran brought in for the lift nobody else will touch: the steel beam that has to thread between two towers in a crosswind. The other operator is nearly as good and charges a fraction of the day rate. For every ordinary lift, pallets of bricks, roof trusses, the day-in, day-out work, you’d be foolish to book the veteran.
GPT-6.1 Sol is the second operator, and this month the gap between the two got very small.
A release that came out of a cancellation
OpenAI launched GPT-6.1 Sol at its DevDay event on September 29, 2026, only a week after GPT-6 Sol. The timing had a backstory. The day before, OpenAI had scrapped the planned launch of GPT-6.1 Astra after internal tests found that model didn’t reliably stay within the scope and permissions of its tasks. Sol 6.1 is a separate, smaller model, and its own safety report looks cleaner: in OpenAI’s test of whether an agent admits that its search tool is broken instead of guessing, GPT-6.1 Sol failed to say so in 2.1% of cases, compared with 4.9% for GPT-6 Sol and 1.5% for GPT-6 Astra.
OpenAI’s pitch is short: near-Astra intelligence for a fifth of the price.
Is it really near Astra?
This is the part that matters, and for once we don’t have to take the vendor’s word for it.
Artificial Analysis, an independent testing firm, ran its benchmarks on launch day. On its Intelligence Index, a combination of ten hard evaluations, GPT-6.1 Sol at max effort scored 52, one point below GPT-6 Astra’s 53 and four above GPT-6 Sol. On its Coding Agent Index, the result is even closer: at max effort Sol sits 2 points behind Astra, and at the second-highest setting, xhigh, it beat Astra by a point for less than 15% of the cost per task.
Cognition, the company behind the Devin coding agent, tested it on FrontierCode 1.1, a benchmark that asks a strict question: would a senior engineer actually merge this change? GPT-6.1 Sol’s best score is 60.4%, essentially the same as GPT-6 Sol (60.7%). The difference is cost. At medium effort it spends $0.31 per task, 81% less than GPT-6 Sol spent at max to reach its best. At low effort it scores 58.1% for just $0.21, up from 50.5% for GPT-6 Sol at the same setting.
OpenAI’s own figures point the same way. On DeepSWE v1.1, complex engineering tasks in real codebases, its chart shows 75.2% at high effort for $0.65 per task. That’s 6.4 points above GPT-6 Sol’s best and a hair above GPT-6 Astra’s best (74.1%), which costs $4.43 per task. One oddity is worth knowing: at xhigh and max, GPT-6.1 Sol actually scores lower, 71.9%. More thinking isn’t always better thinking.
So the honest summary is: roughly Astra-class coding, and the same merge quality as last week’s Sol, at a much lower price per finished job.
Why the cache price is the real headline
A coding agent never reads your project just once. Every time it runs a test, reads the error, and tries again, it re-reads the same files, instructions, and history. Across a long session that can mean the same few hundred thousand tokens going back into the model dozens of times.
That’s what cached input covers: text the model has seen recently in this session. GPT-6.1 Sol charges $0.10 per million cached tokens, 95% off the normal $2 and half of GPT-6 Sol’s cache price. Normal input and output prices stay at $2 and $10 per million, a fifth of GPT-6 Astra’s $10/$50.
Add it up and you get the number that matters most: Artificial Analysis measured about $0.72 per task to run its index with GPT-6.1 Sol, against $3.26 for GPT-6 Astra, $5.98 for Claude Opus 5.5, and $7.60 for Claude Sonnet 5.5. One caution: prompts longer than 272,000 tokens cost double for input and 1.5 times for output, so truly enormous single requests get pricier.
Don’t turn the dial all the way up
GPT-6.1 Sol has five effort settings: low, medium (the default), high, xhigh, and max. It’s tempting to assume max is always best. On Artificial Analysis’s Coding Agent Index, it isn’t: xhigh scored higher than max, and it costs less. For routine fixes, low or medium is often enough, as Cognition’s 58.1%-for-$0.21 result shows. Save xhigh for the problems that really need the extra thinking.
The honest limits
- Better value, not better code. On FrontierCode 1.1, merge quality is flat against GPT-6 Sol and GPT-5.6 Sol. If your problem was that Sol’s patches weren’t good enough, 6.1 doesn’t fix that. It makes the same patches much cheaper.
- Science is still Astra’s job. On OpenAI’s Terminal-Bench Science 0.1, data analysis, simulations, and theorem proving in a terminal, GPT-6.1 Sol scores 57.0% at max effort. That’s more than double GPT-6 Sol’s 27.6%, but well behind GPT-6 Astra (68.1%) and Claude Opus 5.5 (63.3%). It does cost $5.47 a task, against about $23 for either rival.
- Claude scores higher on the broad index. Claude Opus 5.5 (58) and Claude Sonnet 5.5 (56) remain ahead of GPT-6.1 Sol (52) on Artificial Analysis’s Intelligence Index, and although Artificial Analysis measured a 12-point Terminal-Bench 4.0 gain for GPT-6.1 Sol over GPT-6 Sol, Sonnet’s 64% on that test is still higher. Sonnet gets there by burning far more tokens, so Sol wins on cost per task.
- Not in regular Chat. In ChatGPT, GPT-6.1 Sol is available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. Developers use it in the API as
gpt-6.1-sol. OpenAI says a faster “Ultrafast” mode for Codex is coming, but it isn’t live yet.
Who should use it
| If your work looks like… | Reach for | Why |
|---|---|---|
| Bugs, tests, refactors, agents running in CI all day | GPT-6.1 Sol | Near-Astra coding at a fraction of the cost per task |
| Scientific computing, the hardest computer-use jobs | GPT-6 Astra | Still clearly ahead on science terminals |
| Open-ended architecture and judgment calls | Claude Opus 5.5 | Highest on the broad independent index |
| Long terminal sessions where tokens are no object | Claude Sonnet 5.5 | Stronger independent terminal score, higher token use |
The coding verdict
GPT-6.1 Sol doesn’t raise the ceiling. What it does is bring the ceiling within reach of an everyday budget. Independent tests put it a point or two from OpenAI’s flagship at under a quarter of the cost per task, and its cache pricing is built for the way coding agents actually work. Hand it the backlog, set it to xhigh for the hard tickets, and call in Astra or Opus only when it gets stuck.