Skip to content
AI Primer
release

OpenAI launches GPT-6 Astra Ultrafast at up to 8× Standard speed

OpenAI says Ultrafast generates tokens up to eight times faster than Standard in Codex. The tier is available through the API and selected subscriptions at a higher price; GPT-6.1 Sol support is planned.

4 min read
OpenAI launches GPT-6 Astra Ultrafast at up to 8× Standard speed
OpenAI launches GPT-6 Astra Ultrafast at up to 8× Standard speed

TL;DR

  • GPT-6 Astra Ultrafast generates up to 300 tokens per second in Codex, up to 8× Astra Standard; API generation is up to 6× faster, according to OpenAI's launch thread.
  • Developers can use it through the API, while Codex and ChatGPT Work access comes through Pro 500 or Enterprise, per OpenAI's developer thread.
  • The API charges six times Astra Standard's token rates: a pricing-table screenshot lists $60 per million input tokens and $300 per million output tokens for short-context requests.
  • A $500 subscription buys access, not unlimited Ultrafast work: one early buyer reported exhausting his weekly allowance in under 30 minutes.

The API guide exposes Ultrafast as a request setting and recommends persistent WebSockets for tool-heavy agents. One Pro 500 user's two-minute session reportedly consumed 2% of his weekly allowance. The same API guide excludes EU and other non-US regional processing endpoints.

Astra at 300 tokens per second

OpenAI puts Codex generation at up to 8× Astra Standard and 4× Astra Fast. Its launch thread gives the API a separate ceiling of up to 6×; the API guide's headline says up to 8× versus Standard without a matched test setup to reconcile the figures. These are token-generation claims, not end-to-end agent-task timings.

GPT-6.1 Sol Ultrafast is still pending. Its standard model has a 95% cached-input discount, according to thsottiaux's launch post.

The API request flag

The Ultrafast guide keeps the model slug gpt-6-astra and sets service_tier to ultrafast on a Responses request:

HTTP works too. For successive tool calls, OpenAI recommends one persistent WebSocket connection and passing previous_response_id between turns to reduce connection overhead; its developer roundup places Ultrafast among the Codex updates.

Six-times API token pricing

The API pricing page separates Standard, Fast and Ultrafast charges. Astra Ultrafast's posted rates per million tokens are:

  • Short context: $60 input, $6 cached input, $75 cache writes, $300 output.
  • Long context: $120 input, $12 cached input, $150 cache writes, $450 output. The Astra model page says the long-context threshold is more than 272K input tokens and its higher rates apply to the full request.

A product-team DevDay recap describes the trade as 8× speed for 6× cost. At sustained maximum generation speed, that combination could spend about 48× as much on output per wall-clock second, kunchenguid calculated; time spent waiting on tools changes the actual rate.

Pro 500 and the shared allowance

The $500 monthly Pro 500 plan includes Ultrafast in Codex and ChatGPT Work and 25× the Plus usage allowance. The API remains available to developers without that subscription, as OpenAI's plan announcement makes clear.

Included subscription usage drains at 8× the Standard rate on Ultrafast, while purchased credits and eligible Enterprise pay-as-you-go are charged at 6×, according to btibor91's DevDay breakdown. Work and Codex draw from a shared allowance. OpenAI also committed to keeping the five-hour limit removed, leaving the weekly allowance as a constraint.

A weekly limit gone in 30 minutes

After upgrading, one user reported that an Astra Ultrafast prompt running for two minutes used 2% of his weekly allowance. He later posted a Codex screenshot showing 0% left after under 30 minutes; in a follow-up, he said the Sol subagents people had questioned were running on a separate account. The reports describe individual sessions, not a fixed number of prompts included with Pro 500.

API rate limits and residency

The API guide's availability table gives Astra Ultrafast its own default token-per-minute ceilings:

  • Usage tiers 1–3: 500,000 TPM.
  • Tier 4: 1 million TPM.
  • Tier 5: 5 million TPM.

It supports global processing and US data residency, but no EU or other non-US regional processing endpoint. The guide says account teams can arrange higher limits.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 6 threads
TL;DR3 posts
Astra at 300 tokens per second2 posts
The API request flag1 post
Six-times API token pricing3 posts
Pro 500 and the shared allowance2 posts
A weekly limit gone in 30 minutes2 posts
Share on X