GPT-6.1 Sol scores 86.3% on MathArena BrokenArXiv
GPT-6.1 Sol scored 86.3% on MathArena BrokenArXiv, 78.6% on Braintrust problem solving, and third in Code Arena WebDev. A user also reports that it outperformed its mini-SWE result in Codex.

TL;DR
- On MathArena’s 08/2026 BrokenArXiv track, GPT-6.1 Sol scored 86.31%, ranked first of 13 models, and cost $0.94 per run, according to haider1’s MathArena result.
- Braintrust measured 78.6% problem-solving accuracy, up from GPT-6 Sol’s 75.8%, and 77.6% teaching quality, up from 64.1%, in Braintrust’s comparison.
- Code Arena placed GPT-6.1 Sol Max third in WebDev at 1,759 points, 70 points above GPT-6 Sol, at an $8 blended price per million tokens, per Arena’s leaderboard post.
- Coding results depend heavily on the harness: theo’s updated run says Sol performed far better in Codex than mini-swe, while ArtificialAnlys’ reply reports no significant Codex lift across effort levels except low.
- OpenAI lists GPT-6.1 Sol at $2 input, $10 output, and $0.10 cached input per million tokens, with a 1.05 million-token context window, in OpenAIDevs’ launch details.
BrokenArXiv deliberately turns recent arXiv claims into plausible false statements, so it rewards catching a bad premise rather than obediently proving it, according to the BrokenArXiv methodology. MathArena’s GPT-6.1 Sol model page separates the 86.31% August cohort from a 92.10% all-track aggregate. The coding result has its own wrinkle: Theo’s rerun and Artificial Analysis’s report saw different effects from the Codex harness.
The release also filled a gap after TheRundownAI’s launch summary reported that GPT-6.1 Astra would not ship as expected. Sol became OpenAI’s public near-flagship comparison point, with the early evidence split between raw capability, agent harness behavior, and refusal policy.
BrokenArXiv
The headline 86.3% belongs specifically to MathArena’s 08/2026 BrokenArXiv set. The current MathArena results page records the same run as 86.31% ±9.00%, first of 13, at $0.94 and 71,792 output tokens.
- The all-track BrokenArXiv aggregate is 92.10% ±3.58%, first of 12, at $0.44 and 35,600 output tokens, per MathArena’s model page.
- In the supplied comparison, GPT-6 Astra scored 81.94% at $2.26, while Claude Opus 5.5 scored 77.38% at $2.41, according to haider1’s benchmark post.
BrokenArXiv extracts problems from recent papers, perturbs them into statements that are plausible but false, and asks the model to prove them. A correct response recognizes that the statement is wrong as written, which makes the score a measure of mathematical skepticism alongside derivation.
Braintrust
Braintrust’s linked evaluation used 125 problem-solving cases and 50 teaching cases. Problem-solving accuracy moved from 75.8% for GPT-6 Sol to 78.6% for GPT-6.1 Sol, a 2.8-point increase, according to Braintrust’s evaluation.
Teaching quality showed the larger change, rising from 64.1% to 77.6%, a 13.5-point increase, versus Astra’s 82.2%, per Braintrust’s results. The evaluation attributes the gain to better mistake correction and guidance that helps without revealing the answer.
Writing quality stayed roughly flat, with GPT-6.1 Sol near Astra’s result, according to the Braintrust write-up. The improvement is concentrated in how the model leads a user through a problem, rather than in every measured behavior.
Agent benchmarks
OpenAI’s launch results put the model close to Astra across several agent workloads:
- DeepSWE v1.1: 68.8% for GPT-6 Sol to 75.2% for GPT-6.1 Sol at lower effort, a 6.4-point increase, with about 76% lower cost per task, according to OpenAIDevs’ DeepSWE result.
- AutomationBench: GPT-6.1 Sol improved by 4.8 points over GPT-6 Sol at the same setting, while haider1’s AutomationBench result says it beat Opus 5.5 at roughly one-third the cost.
- OSWorld 2.0: GPT-6.1 Sol reached 71.4% versus Astra’s 73.5%, a 2.1-point gap, at roughly one-seventh of Astra’s task cost, per haider1’s OSWorld result.
OpenRouter lists the model at $2 per million input tokens, $10 output, and $0.10 cached input, while its follow-up says Sol matches Astra on DeepSWE and comes within 2.1 points on OSWorld, in OpenRouter’s model listing and OpenRouter’s benchmark follow-up. reach_vb’s DeepSWE result highlighted the same Astra parity, and reach_vb’s rollout reply pointed users to the live rollout.
Token economics
Artificial Analysis scored GPT-6.1 Sol at 52 on its Intelligence Index, up from 48 for GPT-6 Sol and one point below Astra’s 53. Its max-effort cost per task was $0.72, versus $1.05 for GPT-6 Sol and $3.26 for Astra, according to ArtificialAnlys’ cost analysis.
The list price stayed at $2 input and $10 output, but the accounting changed underneath it. Artificial Analysis found 10% to 30% more output tokens than GPT-6 Sol across effort levels, while lower input use and the cache-read discount pushed total task cost down, per ArtificialAnlys’ token breakdown.
The resulting tradeoff is visible in the product surface: OpenAI’s launch post positions Sol as near-Astra quality at one-fifth of Astra’s standard token prices, while OpenRouter also exposes a higher-cost Pro reasoning mode for harder requests, as described by OpenRouter’s Pro-mode note.
Code Arena
Code Arena’s WebDev table ranked GPT-6.1 Sol Max third at 1,759 points. Opus 5.5 Max led at 1,818, Astra Max followed at 1,789, and GPT-6 Sol Max sat seventh at 1,689, according to Arena’s leaderboard.
The same comparison places Sol within 30 points of Astra at 80% lower blended token cost. Arena’s follow-up also reports a 70-point improvement over GPT-6 Sol at the same $8 per million blended price, in Arena’s cost-efficiency comparison.
Harnesses
Theo’s updated Terminal-Bench run found that GPT-6.1 Sol performed “WAY better” in Codex than in mini-swe, the harness used by Artificial Analysis, according to theo’s Codex comparison.
Artificial Analysis replied that its three-repeat test did not show a significant Codex performance bump across effort levels versus its reference harness, apart from low effort, and asked to compare confidence intervals, in ArtificialAnlys’ response. The score therefore belongs to the model-plus-harness combination, with tool loop and evaluation setup capable of changing the ranking.
Workflow reports
The hands-on reports split Sol by task. theo kept Opus 5.5 as the default for implementation, but used GPT-6.1 Sol for code reviews, architecture analysis, computer use, and email management, according to theo’s workflow note.
After a full day, kunchenguid described Sol as capable but slow and difficult to parse, with poor judgment in interactive “firstmate” work. The report categorized it as a solid implementer and reviewer, while kunchenguid’s full-day assessment reserved interactive orchestration for other models.
Higgsfield’s demos show the computer-use surface in more concrete terms. One session used a bottle photo to find a paying brewery client Higgsfield’s bottle demo, while another researched gigs, wrote prompts, compared generated takes, and uploaded 20 short videos through computer use, as shown in Higgsfield’s multitasking demo.
A separate heavy-use report said a day of work moved the weekly subscription meter by barely 2%, according to LLMpsycho’s usage report.
Cybersecurity evaluations
ValsAI’s comparison reports CyberBench falling from 78.0% for GPT-6 Sol to 39.3% for GPT-6.1 Sol. It attributes the decline to stricter cybersecurity filters, records the proof-of-concept portion moving from 0% refusals to 100% refusals, and says 101 of 262 SRE Bench tasks were blocked, according to ValsAI’s benchmark note.
OpenAI’s GPT-6.1 Sol system card reports a different cyber evaluation, ExploitBench, where Sol reached 99.7% at maximum reasoning versus 81.7% for GPT-6 Sol and 100% for Astra. CyberBench and ExploitBench test different tasks, so the two sets of numbers describe a capability and filtering split rather than a single clean regression line.