GPT-6.1 Sol adds Ultrafast mode at up to 8× Standard speed
OpenAI says GPT-6.1 Sol Ultrafast runs up to eight times faster than Standard. API pricing is $12/$60 per million input/output tokens, while Codex and ChatGPT Work access is limited to eligible plans.

TL;DR
- GPT-6.1 Sol Ultrafast is rolling out across the API, Codex, and ChatGPT Work, with OpenAIDevs claiming up to 8x Standard speed.
- API pricing rises to $12/M input and $60/M output, according to OpenAI's pricing announcement, six times Sol Standard's rates.
- Bundled Codex and Work access requires Pro 500, eligible usage-based Enterprise, or credit-based Edu plans, with Enterprise admin approval, per OpenAI's eligibility announcement.
- Faster inference leaves other bottlenecks intact: most of MatthewBerman's tasks saw little benefit because the surrounding workflow remained slow.
The full price sheet puts long-context Ultrafast output at $90/M and short-context cached reads at $0.60/M. Vercel's launch notes describe EU-pinned requests falling back to Standard, despite OpenAI's rollout announcement promising EU residency.
What shipped
- API: The documentation captured in rohanpaul_ai's post makes Sol Ultrafast available to all API users, subject to separate rate limits.
- Request configuration: The model remains
gpt-6.1-sol; requests selectservice_tier: "ultrafast"in the same documentation screenshot. - Codex and ChatGPT Work: Access covers Pro 500, eligible usage-based Enterprise, and credit-based Edu plans, according to OpenAI's plan details.
- Enterprise administration: Admins must enable access, OpenAI says in the same announcement.
The Pro 500 restriction applies to bundled product access. API access has its own usage-based billing and eligibility.
API pricing
Sol Ultrafast carries an Astra-sized bill: 6x Sol Standard's token rates, as WesRoth calculated, and 20% above Astra Standard in the pricing comparison below.
OpenAI's pricing table extends the sixfold premium to caching and long context. All figures below are dollars per million tokens, comparing Standard with Ultrafast:
- Short-context cached input: $0.10 → $0.60, +500%.
- Short-context cache writes: $2.50 → $15, +500%.
- Long-context input: $4 → $24, +500%.
- Long-context output: $15 → $90, +500%.
The API changelog defines Sol's short-context pricing band as prompts with up to 272K input tokens.
Benchmarks and evaluation
OpenAI's published capability comparisons come from the original GPT-6.1 Sol release. Those results compare base models; the Ultrafast rollout adds a generation-speed claim.
First-party
- Token generation: Sol Standard → Sol Ultrafast, 1x → up to 8x, up to +700%, according to OpenAI's rollout announcement.
- DeepSWE v1.1: GPT-6 Sol → GPT-6.1 Sol, +6.4 percentage points at lower reasoning effort, in OpenAI's original evaluations.
- AutomationBench: GPT-6 Sol → GPT-6.1 Sol at medium effort, +4.8 percentage points, in the same evaluations.
- Answers containing factual errors, low effort: 11.4% → 7.7%, −3.7 percentage points, on deliberately difficult, user-flagged conversations in OpenAI's factuality evaluation.
Third-party evaluators
Customer-reported
Ultrafast produced better results on some evaluations, pvncher reported, describing the gains as task-specific without supplying scores or a matched Standard-versus-Ultrafast comparison.
API requests and latency limitations
OpenAI originally scoped the eightfold claim to token generation in Codex in Sol's model announcement. The Ultrafast rollout announcement supplies no p50/p95 latency figures or completed-agent timings.
The request shape shown in the documentation screenshot selects the tier explicitly:
OpenAI's Ultrafast guide describes the transport mechanics:
- Persistent WebSockets reduce connection overhead across frequent tool calls.
- Successive turns can reuse the connection and pass
previous_response_id. - HTTP requests remain supported through the SDK.
Ultrafast has separate limits from Standard and Fast in the captured documentation. The guide's numerical TPM table explicitly covers Astra, so it does not establish Sol's Ultrafast quota.
Instant steering
Codex paired the rollout with “instant” steering, allowing the model to react faster to corrections made while work is in progress.
Vibe Check
- UI exploration: Rapid frontend iterations are where dkundel finds Ultrafast useful, followed by a handoff to regular speed once the direction is settled.
- Voice and computer use: Working on a slide deck with voice enabled and the output fullscreen felt substantially different to pvncher, who described talking through changes while watching the result.
- Surrounding infrastructure: Most tasks did not benefit for MatthewBerman, because everything around inference was still slow.
The access gate drew its own reaction. Users first encountered the Pro 500 requirement in the third announcement post, which kimmonismus criticized as a buried restriction.
OpenRouter and AI Gateway
OpenRouter announced live support for Sol Ultrafast.
Its Sol model page exposes the model as openai/gpt-6.1-sol.
Vercel AI Gateway accepts service_tier: "ultrafast" per request, including over WebSocket, according to vercel_dev. The AI SDK example expresses that option as providerOptions.openai.serviceTier and restricts routing to OpenAI with gateway.only: ['openai'].
EU residency and gateway fallback
OpenAI's regional rollout announcement covers three offerings:
- GPT-6.1 Sol Ultrafast: US and EU data residency across supported regions.
- GPT-6.1 Sol Fast: Newly added EU data residency.
- GPT-6 Luna Fast: Newly added EU data residency.
Vercel's same-day changelog describes narrower gateway behavior: Ultrafast supports US and global processing, while requests pinned to unsupported regions, including the EU, run at the Standard default tier. Billing follows the tier actually served, so those fallbacks incur Standard rates.
The Astra-focused availability section of OpenAI's Ultrafast guide also lists US residency and global processing only.
Regional processing carries another price adjustment: OpenAI's pricing page adds a 10% uplift for eligible models released from March 5, 2026 onward. Where that uplift applies, Sol Ultrafast's headline short-context input/output rates become $13.20/$66 per million tokens.