OpenAI launches GPT-6 Astra Ultrafast at up to 8× Standard speed
OpenAI says Ultrafast generates tokens up to eight times faster than Standard in Codex. The tier is available through the API and selected subscriptions at a higher price; GPT-6.1 Sol support is planned.

TL;DR
- GPT-6 Astra Ultrafast generates up to 300 tokens per second in Codex, up to 8× Astra Standard; API generation is up to 6× faster, according to OpenAI's launch thread.
- Developers can use it through the API, while Codex and ChatGPT Work access comes through Pro 500 or Enterprise, per OpenAI's developer thread.
- The API charges six times Astra Standard's token rates: a pricing-table screenshot lists $60 per million input tokens and $300 per million output tokens for short-context requests.
- A $500 subscription buys access, not unlimited Ultrafast work: one early buyer reported exhausting his weekly allowance in under 30 minutes.
The API guide exposes Ultrafast as a request setting and recommends persistent WebSockets for tool-heavy agents. One Pro 500 user's two-minute session reportedly consumed 2% of his weekly allowance. The same API guide excludes EU and other non-US regional processing endpoints.
Astra at 300 tokens per second
OpenAI puts Codex generation at up to 8× Astra Standard and 4× Astra Fast. Its launch thread gives the API a separate ceiling of up to 6×; the API guide's headline says up to 8× versus Standard without a matched test setup to reconcile the figures. These are token-generation claims, not end-to-end agent-task timings.
GPT-6.1 Sol Ultrafast is still pending. Its standard model has a 95% cached-input discount, according to thsottiaux's launch post.
The API request flag
The Ultrafast guide keeps the model slug gpt-6-astra and sets service_tier to ultrafast on a Responses request:
HTTP works too. For successive tool calls, OpenAI recommends one persistent WebSocket connection and passing previous_response_id between turns to reduce connection overhead; its developer roundup places Ultrafast among the Codex updates.
Six-times API token pricing
The API pricing page separates Standard, Fast and Ultrafast charges. Astra Ultrafast's posted rates per million tokens are:
- Short context: $60 input, $6 cached input, $75 cache writes, $300 output.
- Long context: $120 input, $12 cached input, $150 cache writes, $450 output. The Astra model page says the long-context threshold is more than 272K input tokens and its higher rates apply to the full request.
A product-team DevDay recap describes the trade as 8× speed for 6× cost. At sustained maximum generation speed, that combination could spend about 48× as much on output per wall-clock second, kunchenguid calculated; time spent waiting on tools changes the actual rate.
Pro 500 and the shared allowance
The $500 monthly Pro 500 plan includes Ultrafast in Codex and ChatGPT Work and 25× the Plus usage allowance. The API remains available to developers without that subscription, as OpenAI's plan announcement makes clear.
Included subscription usage drains at 8× the Standard rate on Ultrafast, while purchased credits and eligible Enterprise pay-as-you-go are charged at 6×, according to btibor91's DevDay breakdown. Work and Codex draw from a shared allowance. OpenAI also committed to keeping the five-hour limit removed, leaving the weekly allowance as a constraint.
A weekly limit gone in 30 minutes
After upgrading, one user reported that an Astra Ultrafast prompt running for two minutes used 2% of his weekly allowance. He later posted a Codex screenshot showing 0% left after under 30 minutes; in a follow-up, he said the Sol subagents people had questioned were running on a separate account. The reports describe individual sessions, not a fixed number of prompts included with Pro 500.
API rate limits and residency
The API guide's availability table gives Astra Ultrafast its own default token-per-minute ceilings:
- Usage tiers 1–3: 500,000 TPM.
- Tier 4: 1 million TPM.
- Tier 5: 5 million TPM.
It supports global processing and US data residency, but no EU or other non-US regional processing endpoint. The guide says account teams can arrange higher limits.