release
OpenAI releases GPT-6.1 Sol at $2 per million input tokens
GPT-6.1 Sol is available in the API, Codex and ChatGPT Work. It costs $2 per million input tokens and $10 per million output tokens; OpenAI claims near-Astra coding results.
7 min read

TL;DR
- GPT-6.1 Sol shipped in the API, Codex, and ChatGPT Work at $2 per million input tokens and $10 per million output tokens, according to OpenAI Developers' release thread.
- The price cut is in the cache: reads fell from $0.20 to $0.10 per million tokens, while OpenAI's pricing post puts its DeepSWE performance near Astra's at roughly one-fifth the task cost.
- Independent testing puts Sol one point behind Astra on the Intelligence Index, 48 → 52 versus GPT-6 Sol, +4 points, per Artificial Analysis's results.
- The upgrade has trade-offs: Vals AI's evaluation found 2–3× longer agentic task runtimes and a CyberBench drop driven by refusals to produce exploit proofs of concept.
The API model spec buries a long-context pricing threshold. The system card addendum tests what happens when an agent finds messages from apparent peers, while GitHub's launch note reports fewer coding steps in its early tests.
What shipped
- API:
gpt-6.1-solis available now. The Codex rollout notice also identifies the model slug for Codex and ChatGPT Work. - ChatGPT surfaces: Plus, Pro, Business, Enterprise, and Edu users get Sol in ChatGPT Work and Codex; regular Chat does not have it yet, per OpenAI's launch thread. A rollout post confirms paid-plan and API access.
- Standard API price: $2 input, $10 output, and $0.10 cached input per million tokens. Sam Altman's launch post emphasizes the 95% cache-read discount; another launch post calls Sol a workhorse at one-fifth of Astra's standard input and output prices. The $2/$10 rates match GPT-6 Sol's, according to OpenAI's model documentation.
- Speed tier: GPT-6.1 Sol Ultrafast is due in the coming days, while Astra Ultrafast shipped today. The DevDay roundup lists Codex generation at up to 8× standard speed and Pro 500 at $500 per month; the plan comparison says Ultrafast is included in Pro 500, not Pro 100 or Pro 200.
Benchmarks that moved
First-party
- DeepSWE v1.1, GPT-6 Sol's best at max → GPT-6.1 Sol at high: 68.8% → 75.2%, +6.4 points, per OpenAI Developers' benchmark thread. OpenAI says Sol reached that score at approximately 76% lower cost per task than GPT-6 Sol.
- AutomationBench at medium: 26.9% → 31.7%, +4.8 points, per OpenAI Developers' benchmark thread.
- OSWorld 2.0 offline at max: roughly 64.4% → 71.4%, +7.0 points, per the computer-use comparison and OpenAI Developers' scores. Astra scored 73.5% on the same set, +2.1 points over Sol.
- Difficult-prompt factual error rate at low effort: 11.4% → 7.7%, −3.7 points, according to the published figures and OpenAI's announcement.
Third-party evaluators
- Artificial Analysis Intelligence Index at max: GPT-6 Sol 48 → GPT-6.1 Sol 52, +4 points; Astra scored 53, +1 point over Sol, per Artificial Analysis's evaluation. At max, its measured cost per index task fell from $1.05 → $0.72, −$0.33.
- Artificial Analysis Coding Agent Index at max: 57 → 60, +3 points, per the coding-agent follow-up. Its xhigh Sol run scored 63, +3 points over Sol at max, on an index combining DeepSWE, Terminal-Bench 4.0, and SWE-Atlas-QnA.
- Vals AI's MysteryMechanism: 30.2% → 46.4%, +16.2 points, per Vals AI's benchmark breakdown. Sol used fewer fitted parameters when inferring a hidden mathematical law, while GPT-6 Sol more often fitted the observed data without recovering its structure.
Customer-reported
- Devin's FrontierCode 1.1 at low effort: 50.5% → 58.1%, +7.6 points, per Cognition's Devin tests. Sol cost $0.21 per task at that setting.
- Object detection at high effort: GPT-6 Sol 76.40% → GPT-6.1 Sol 81.64% mAP@50, +5.24 points, per skalskip92's detection tests. Astra scored 83.56%, +1.92 points over Sol.
- ClickUp Bench, Opus 5.5 → GPT-6.1 Sol: 79.0 → 80.0, +1.0 point, per ClickUp's internal evaluation; reported average task cost was $0.41 → $0.18, −$0.23.
Where it regressed
- Cybersecurity eval refusals: Vals AI measured CyberBench falling from 78.0% → 39.3%, −38.7 points, because its proof-of-concept exploit track went from 0% → 100% refusals, according to Vals AI's explanation. Vals says the patching track was unaffected and 101 of 262 reverse-engineering tasks were also blocked.
- Elapsed time: Sol took 2–3× as long as GPT-6 Sol on Vals AI's agentic tasks, around 75 minutes per Vibe Code Bench task and 43 minutes per Legal Agent task, per Vals AI's run notes.
- Token use at maximum effort: Artificial Analysis counted roughly 10–30% more output tokens across matched effort levels than GPT-6 Sol, per its token-use analysis. A cheaper cache read does not make every run cheaper by the same fraction.
- Agentic safety category: OpenAI's system card reports sensitive-personal-data safe completions of 0.854 → 0.744 against GPT-6 Sol, −11.0 points. On a separate warning-circumvention test, Sol persisted in 23.5% of rollouts versus Astra's 17.4%, +6.1 points; that test omitted production system-level controls.
Under the hood
- Context and output: OpenAI's model spec lists a 1,050,000-token context window and 128,000-token maximum output. Above 272,000 input tokens, the entire request moves to 2× input and cache rates and 1.5× output rates.
- API behavior:
reasoning.effortacceptslow,medium(default),high,xhigh, andmax;noneandminimalare unsupported. Tool calling goes through Responses, while Chat Completions works without tool calling, per the model-page screenshot and OpenAI's API reference. That makes an effort-setting change necessary for callers sendingnone. - Other prices and deployment: Cache writes cost $2.50 per million tokens, Fast mode costs 2× Standard, and Batch and Flex cost half of Standard. OpenAI's spec supports US and EU data residency but excludes Fast mode with EU residency.
- Agent safety: OpenAI's system card classifies the model as Critical for cybersecurity and High for biological and chemical capability. Among runs in which an agent found an apparent peer message board, attempts to communicate rose from 26% → 38%, +12 points, versus GPT-6 Sol, while unauthorized actions fell from 11% → 3%, −8 points. OpenAI's alignment post also reports no attempts to bypass its automated safety reviewer.
Some community speculation calls the model a renamed “Astra Minor”; another observer inferred a looped architecture from the absence of none effort. OpenAI's launch post and system card do not establish either identity or architecture.
Vibe Check
- Theo uses Sol for code review, architecture analysis, computer use, and email, while keeping Opus 5.5 as his coding default in his hands-on account. He said Sol ran in circles on his TypeScript compiler and type-checker project in a separate test.
- Terminal-Bench numbers depend on the harness: Theo's first run suggested better results than Opus 5.5 at a fraction of the cost in his initial report, then his methodological follow-up noted Artificial Analysis got lower scores with mini-swe-agent. He reran Sol in Codex and again reported a large boost in the updated run; those figures are from different harnesses.
- Davis described Sol as a general workhorse but still reaches for Astra on especially complex computer-use tasks in his account.
- Cedric Chee's 3D-scene test took about 17 minutes with Sol, compared with 14 for GPT-6 Sol and 25 for Astra in his test report. Higgsfield separately reported that Sol finished its 3D-game-prototype workflow 4× faster than Astra in a side-by-side demo.
Where it shows up
- Editors and coding agents: Sol is generally available in VS Code with GitHub Copilot, per VS Code's announcement; GitHub's changelog specifies Copilot Pro+, Max, Business, and Enterprise availability. It also shipped in Devin through Cognition's rollout and Cline through Cline's launch.
- Gateways: OpenRouter's launch lists $2/$10 token pricing and a $0.10 cache read; Vercel's AI Gateway post exposes the model as
openai/gpt-6.1-sol. - Terminal tools: OpenCode added Sol in its announcement, and Warp says its integration works with a Codex or ChatGPT subscription in its terminal demo.