Skip to content
AI Primer
release

GPT-6 Sol reportedly launches in Codex, Work, and the API

Posts report that GPT-6 Sol is live in Codex, Work, and the API alongside the lighter GPT-6 Luna. Reported API prices are roughly half those of the prior Sol models.

6 min read
GPT-6 Sol reportedly launches in Codex, Work, and the API
GPT-6 Sol reportedly launches in Codex, Work, and the API

TL;DR

  • Sol is positioned for harder everyday work and Luna for lighter, cheaper tasks, as ozansihay's launch-day summary described them; OpenAI says both now ship in Work, Codex, and the API, a rollout minchoi's launch post also reported.
  • API spend is the release's clearest change: Sol drops from $4/$20 to $2/$10 per million input/output tokens, while minchoi's reply framed the move as direct cost competition.
  • Long autonomous coding remains the early split: danshipper's test report found Sol improved on GPT-5.6, but gave Opus 5.5 the higher ceiling on extended builds.
  • Creative one-prompt comparisons already include a WebGL-heavy studio site, where viktoroddy's WebGL comparison gave both models the same brief: an interactive sphere, oversized type, and smooth scrolling.

OpenAI pairs the price cut with a 90% discount on cached input reads in its launch post. Artificial Analysis found Sol's lower hallucination rate coincided with answering 83% rather than 99% of its questions, while Devin said Sol reached the same FrontierCode score while reading about 17% less context. Higgsfield's higgsfield_ai's anatomy demo turns the model matchup into a take-apart 3D eye.

What shipped

  • GPT-6 Sol and GPT-6 Luna are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu accounts, and in the API as gpt-6-sol and gpt-6-luna, according to OpenAI's availability note.
  • Free and Go accounts can use Luna in the desktop app; OpenAI says neither model is yet available in Chat.
  • Sol API input/output: $4/$20 → $2/$10 per million tokens, -50%, in OpenAI's pricing table.
  • Luna API input/output: $0.20/$1.20 → $0.10/$0.50 per million tokens, -50% input and -58.3% output, in the same table.
  • OpenAI says access is rolling out gradually during the day to protect service stability.

Benchmarks that moved

First-party

  • AutomationBench 1.0.6: Claude Opus 5 at max effort 26.9% → GPT-6 Sol at xhigh effort 33.2%, +6.3 points, per OpenAI's table.
  • OSWorld 2.0 offline: Claude Opus 5 at medium effort 60.3% → GPT-6 Sol at xhigh effort 60.5%, +0.2 points, in OpenAI's computer-use comparison.

Third-party evaluators

Customer-reported

  • Every's six-task Hands check: GPT-6 Astra 17/18 attempts → an early Sol preview 17/18, 0 points, in danshipper's Vibe Check.

Where it regressed

Artificial Analysis found Sol's Knowledge Agent accuracy fell from 59% to 54%, -5 points, even as its hallucination rate dropped from 92% to 60%, -32 points. The evaluator attributes much of that shift to Sol declining more questions, 83% attempted versus 99% for GPT-5.6 Sol; OpenAI's own factuality evaluation instead reports roughly half as many mistakes on flagged real-world conversations.

The same independent evaluation puts Sol about 100 Elo lower on GDPval-AA v2.1, Luna about 75 Elo lower there and about 45 Elo lower on AA-Briefcase. Its reviewers traced the knowledge-work misses to reduced presentation quality and deliverables that omitted rubric elements.

Luna also falls from 43 to 41 on the Artificial Analysis Coding Agent Index, -2 points, with lower SWE-Atlas-QnA and DeepSWE results than GPT-5.6 Luna.

Every CEO and cofounder Dan Shipper separately reported that Codex's security classifier repeatedly halted work the team had already authorized, making long runs harder to leave unattended in danshipper's test report. OpenAI's launch post does not explain that classifier behavior.

Under the hood

  • Cache economics: OpenAI says in its launch post that cached input reads receive a 90% discount. Artificial Analysis reports the corresponding cache writes carry a 25% premium.
  • Cache preservation: changing reasoning effort or turning tools on and off now preserves earlier context for reuse, according to OpenAI; developers can also set explicit cache breakpoints and inspect misses in its Prompt Caching Dashboard.
  • Response style: OpenAI says Sol and Luna inherit Astra's shorter, more direct technical communication, with fewer low-value details and odd phrasing.
  • Context window: OpenAI's public launch announcement gives no canonical context-window figure, despite specifying pricing, model IDs, cache behavior, and access surfaces.

Vibe Check

Dan Shipper, CEO and cofounder of Every, ran writing, computer-use, and coding tests tied to his company's work rather than a general leaderboard in Every's full Vibe Check.

  • Shipper said in danshipper's Vibe Check that Sol's writing came close to Astra's on a paragraph task, with cleaner, more front-loaded prose.
  • The same runs gave Sol the advantage over GPT-5.6 on coding, but put Opus 5.5 ahead for long autonomous builds and complex visual projects.
  • petergyang's early take similarly called Sol better than GPT-5.6 while naming Opus 5.5 his preferred model over Fable and Astra.

Creative builds

  • A creative-studio landing page with an interactive WebGL sphere, oversized typography, and scroll animation.

Claude Opus 5.5 and GPT-6 Sol build the same WebGL studio-site brief

  • An Unreal Engine 3D game environment and playable walkthrough.

Two model-driven Unreal game builds

  • A rotatable, layer-by-layer 3D human eye that turns an anatomy diagram into an explorable object.

Interactive anatomy site made with Higgsfield

  • A 3D-printable prosthetic hand, plus a site containing a manipulable model, CAD files, and assembly instructions.

Prosthetic-hand builds from the Sol and Opus comparison

  • Two commercial motion-design directions for the same product.

Side-by-side product motion design

  • A JavaScript animation that wakes a digital mosaic.

Two coded mosaic animations

Where it shows up

  • Linear coding sessions now offer Sol, Luna, and Claude Opus 5.5, as Linear said in linear's announcement.
  • GitHub Copilot describes Sol as its balanced choice for interactive and agentic coding, and Luna as the low-cost option for smaller, faster tasks.
  • Devin added both models to Devin Desktop and Devin CLI on launch day.
  • Microsoft Foundry made the two models generally available for production agents, while AWS announced general availability on Bedrock.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 6 threads
TL;DR4 posts
What shipped3 posts
Benchmarks that moved1 post
Under the hood1 post
Vibe Check1 post
Creative builds6 posts
Share on X