Skip to content
AI Primer
update

OpenAI adds capacity after GPT-6.1 Sol overloads ChatGPT

OpenAI added capacity after GPT-6.1 Sol became its most demanded model and caused heavy load in ChatGPT and Codex. OpenAI says the added capacity should nearly double speeds from the prior day.

5 min read
OpenAI adds capacity after GPT-6.1 Sol overloads ChatGPT
OpenAI adds capacity after GPT-6.1 Sol overloads ChatGPT

TL;DR

  • OpenAI added serving capacity after GPT-6.1 Sol became its most demanded model across API and subscriptions, with speed expected to nearly double from the prior day thsottiaux's capacity update.
  • The rollout produced both hard capacity errors and slow responses, including a Codex screen showing “Selected model is at capacity” capacity error screenshot and a report of slow Sol Ultra LearnOpenCV's slow-model report.
  • Sol paired near-Astra positioning with $2 input and $10 output per million tokens, while Artificial Analysis placed it one point below Astra at less than a quarter of Astra's cost per task Artificial Analysis's comparison.
  • The same release cycle reset subscription economics, moving the published Pro 200 multiplier from 20x to 10x while existing users temporarily kept the old multiplier thsottiaux's plan update.

OpenAI's launch post put Sol in ChatGPT Work and Codex, not regular Chat, while the API model page exposes a 1.05 million-token context window, tool support through Responses, and $0.10 cached input. Artificial Analysis put the model one point below Astra at under a quarter of its task cost. An OpenAI community thread records one Plus user exhausting an allowance after roughly nine minutes of visible Codex work.

The overload

GPT-6.1 Sol arrived on September 29 as OpenAI's lower-cost model for complex coding, computer use, and professional work. The official announcement described near-Astra performance at one-fifth of Astra's standard API input and output prices.

The first distribution push was broad. An early post called Sol live in the API at near-Astra performance for less money early Sol announcement, while a contemporaneous daily brief listed it in Work, Codex, and the API at $2/$10, with cached input at $0.10 the daily brief. Artificial Analysis later measured a four-point gain over GPT-6 Sol and a one-point gap to Astra, with a max-effort cost per task of $0.72 versus Astra's $3.26 Artificial Analysis's comparison.

By the next day, thsottiaux said Sol was “our most demanded model pretty much ever” across API and subscriptions. The same update said ChatGPT and Codex were under heavy load, that more capacity had come online, and that speed should approach twice the previous day's level.

ChatGPT Work and Codex

The rollout surface is narrower than the word “ChatGPT” suggests. OpenAI's availability details list Plus, Pro, Business, Enterprise, and Edu access in ChatGPT Work and Codex, plus API access through gpt-6.1-sol; the launch page says Sol was not yet available in regular Chat availability details.

That makes Work and Codex the clearest consumer-facing locations for the overload. A Codex screenshot showed a failed command followed by the message “Selected model is at capacity. Please try a different model” capacity error screenshot.

The API had its own demand surge. thsottiaux explicitly grouped API and subscriptions together when describing Sol's demand thsottiaux's demand report, while the public API page lists tiered limits from 500 RPM and 500,000 TPM at Tier 1 to 15,000 RPM and 40 million TPM at Tier 5 API rate-limit documentation.

Serving symptoms

Users reported latency alongside outright rejection. LearnOpenCV asked whether anyone else was seeing “really really slow” GPT-6.1 Sol Ultra LearnOpenCV's slow-model report, and davis7 said token throughput was rough while running five agents, although the lower token count still produced faster overall work davis7's multi-agent report.

A separate OpenAI community thread mixed several signals: users described Codex allowances draining after a few minutes, while another report logged 3.48 billion input tokens and 11 million output tokens in a session that also ran slowly Codex rate-limit reports. The latter author later identified a stale instruction and unusually wordy subagents as possible contributors, so the thread does not isolate serving capacity as the cause.

Model behavior also shaped the first-day experience. After a full day with Sol, kunchenguid described it as a solid implementer and reviewer but a poor interactive “firstmate,” citing visible slowness and weak judgment on next-step decisions kunchenguid's hands-on review.

OpenAI's capacity update did not specify queue depth, error rates, GPU count, or regional breakdown. It described a near-term speed improvement after capacity came online, but did not publish a post-incident measurement.

Quota math

The capacity problem arrived alongside a subscription reset. thsottiaux published new relative allowances of Plus 1x, Pro 100 5x, and Pro 200 10x, while saying existing Pro 200 subscribers would temporarily retain the 20x multiplier and receive additional credits thsottiaux's plan update.

Users described the change as a halving of the $200 plan's effective usage Cedric Chee's usage complaint. thdxr offered a different explanation, describing subscriptions as a shared compute pool whose allocation has to be balanced against higher-margin API demand thdxr's compute explanation.

Those are two separate claims in the public record. The quota posts document a change in plan economics; the capacity update documents overloaded service. Neither establishes that the quota change caused the serving incident, though both appeared during the same Sol rollout.

Ultrafast

OpenAI's response included a separate speed tier. The DevDay roundup described Astra Ultrafast as available, Sol Ultrafast as forthcoming, and speed targets of up to 8x in Codex and 6x in the API DevDay roundup. The launch materials described Sol Ultrafast as arriving in the following days with up to eight times faster token generation in Codex Sol speed details.

The API model page currently documents Fast mode as twice the standard price and separately says it is unavailable with EU data residency GPT-6.1 Sol API documentation. In an early hands-on report, martinbowling said Ultrafast was indeed fast and appeared to take fewer side paths in its visible reasoning martinbowling's Ultrafast report.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 5 threads
TL;DR2 posts
The overload3 posts
Serving symptoms2 posts
Quota math2 posts
Ultrafast2 posts
Share on X