Pricing & limits
Pricing changes, quotas, rate limits, and plan/tier changes for AI dev tools and APIs.
Stories
Filter storiesFLUX 3 Image supports bounding-box layouts, targeted edits, native 4K output, and ten reference images. Commercial weights are available, open weights are planned, and API use is half-price through October 8.
Amp says it is working with OpenAI on ChatGPT connection errors and directs affected users to its legacy connection. Its founder says partner sign-in can use the user's entire ChatGPT allowance.
Google is rolling out Gemini 4 Argon through Fairwind to government users, vetted cyber defenders, and trusted testers. The model supports up to 1M output tokens and costs $2/M input and $10/M output.
GPT-6.1 Sol is available in the API, Codex and ChatGPT Work. It costs $2 per million input tokens and $10 per million output tokens; OpenAI claims near-Astra coding results.
Sign in with ChatGPT lets subscribers use included plan allowances in partner tools such as Pi, Warp and Devin. Devin documents quota controls, while Amp says higher-capacity use carries separate charges.
OpenAI says Ultrafast generates tokens up to eight times faster than Standard in Codex. The tier is available through the API and selected subscriptions at a higher price; GPT-6.1 Sol support is planned.
Sonnet 5.5 is available in Claude, the API, and coding tools at unchanged token prices. Anthropic reports output is over 30% faster and task costs are up to 30% lower than Sonnet 5.
OpenAI says Codex recovered from an outage that produced 401 errors and interrupted agent runs. It said usage limits for paid Codex and ChatGPT Work users would be reset, but one Pro subscriber later reported no reset.
Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens, below Opus 5 pricing. Cached-input reads now cost $0.20 per million tokens, and Claude Opus 5.5 is the default in Claude Code.
xAI released Grok 4.7 through Grok Build, APIs, Cursor, and other gateways. Early evaluations report stronger coding and knowledge-work results than Grok 4.6, with mixed results across individual coding benchmarks.
Amp says its coding agent is free when users supply their own compute, model subscription, or API key. The rollout includes routing support for external providers such as Ollama Cloud, OpenRouter, and custom OpenAI-compatible URLs.
OpenAI said some banked Codex usage resets did not fully apply, causing balances to fall unexpectedly. Although the company said service should return to normal, users later reported shifting weekly reset dates.
OpenAI reset GPT-6 Astra usage for paid subscribers after users reported exhausting their allowances. An OpenAI employee said there is no fixed reset schedule.
Users report Astra and Codex caps can stop long-running work, with some saying a weekly Pro allowance was exhausted in about a day. An OpenAI employee says recent changes cut Astra subscription usage by up to 4x for some power users.
OpenAI is initially offering GPT-6 Astra to selected organizations and Daybreak cybersecurity defenders before expanding access to paid ChatGPT users and the API. Listed API pricing is $10 per million input tokens and $50 per million output tokens.
Google released Gemini 3.8 Flash for the Gemini API and Google product surfaces. Input and output pricing remains $0.75 and $3.75 per million tokens, respectively.
Alibaba released Qwen3.8-Max-0902 through QwenCloud with a 1M-token context window and 2.4T parameters. The company prices input at $2 per million tokens and says the model leads Code Arena's WebDev leaderboard.
Anthropic says Claude Fable 5.1 is available across Claude and coding platforms for long-running agent work. It reports 52.6% on Terminal-Bench-Science and API cache-read pricing of $0.25 per million tokens.
Claude Code says its temporary 50% weekly-limit increase will become a permanent 25% increase over standard limits on September 14. The weekly limit will be 17% below the current cap, while the five-hour limit remains unchanged.
OpenAI says it will end Cursor's direct access to its models on November 12 after SpaceX acquired Cursor. Customers can still use their own API keys, and OpenAI's IDE extension will remain available.
Z.ai identified the formerly anonymous Ox Alpha as GLM-5.3-Flash and released it under an MIT license. The native multimodal model has 320B total parameters, 18B active parameters, and a 1M-token context window.
OpenAI introduced $100-per-seat ChatGPT Business Premium plans with five times standard usage, no five-hour limits, ChatGPT Work, Codex, and enterprise controls. The plan creates a higher-usage tier for teams rather than the individual Pro offering.
GPT-5.6 Sol now costs $4 per million input tokens and $10 per million output tokens. Benchmark comparisons place Sol at 72.7% on DeepSWE for $6.47 per task, while OpenAI and AWS report lower successful-task costs for Terra in Kiro.
WAN 3.0 is now available through Replicate, OpenRouter, and ComfyUI for text-, image-, and reference-driven video generation up to 1080p. Pika says the model supports up to 20 references and 30-second output, while quality and price comparisons remain vendor-reported.
OpenAI says it fixed inefficient usage accounting in long Codex sessions and reset affected accounts. Some users report that business accounts or active sessions did not receive the reset.
OpenAI cut GPT-5.6 Sol API input and output rates from $5 and $30 to $4 and $20 per million tokens for three months. Subscription usage remains unchanged.
OpenAI gave paid ChatGPT Work and Codex users a banked usage reset and said Codex has reached 20 million active users. The company is investigating reports that lower cache-hit rates are causing usage limits to drain faster.
OpenRouter is testing Ox Alpha, a free stealth model with a 1M-token context window. The model accepts text, image, and video inputs and is offered with zero data retention during the test.
Nous Research launched Hermes Agent with managed remote computers, provider and model choice, and local or hosted execution. Hermes Cloud idle instances start at 3 cents per day, according to the company.
Z.ai released GLM-5.3 through an API with OpenAI- and Anthropic-compatible interfaces. Input and output pricing remains $1.40 and $2.80 per million tokens, matching GLM-5.2.
OpenCode revised its Go plan to include $30 in usage for a $10 subscription. The change follows a provider price increase and the company’s effort to secure lower-cost capacity.
OpenCode says it revised Go limits after DeepSeek raised prices. Its operator is testing hosting configurations intended to bring DeepSeek service closer to its prior price point.
Codex now lets ChatGPT subscribers enable a 1 million-token context window for GPT-5.6 Sol. Automatic compaction starts at 900,000 tokens, and tokens beyond the default window count double against limits.
OpenCode’s thdxr said Go users spent $1.14 per day on DeepSeek V4 Flash last week. Wafer added a fast OpenRouter route, while Nous extended a 90% discount for the 0731 model.
Guillermo Rauch listed shipped safeguards including spend caps, anomaly alerts, recursion protection, billing APIs and DDoS mitigation. The post followed reports of agents looping until queues timed out.
ComfyUI says Seedance 2.5 is live via Partner Nodes with 30-second runs, up to 50 references, timeline shot control, editing, and multilingual lip sync. Fal, Pika, Venice, and other tools also added access.
Cursor users say AI IDE agent runs are hard to audit and hard to constrain. Reddit threads cite hidden model routing, cache charges, destructive SQL migrations, rules folders, runbooks, and context-trimming pipelines.
Cline tripled its free quota while Nous discounted Flash 0731 by 90% for a week and OpenHands offered free cloud use. OpenCode reported 8T Flash tokens on Aug. 1.
Claude Code users reported Fable 5 and Opus 5 sessions exhausting five-hour usage windows in about 6–10 minutes after automated tool calls. The reports tie the failures to agent loops and rate-limit economics, including one Reddit claim of 10.26M tokens, 15 calls, and no edits.
OpenAI said GPT-5.6 Luna pricing fell 80%, while Terra fell 20%. Codex users recommended max reasoning for linting, tests, and dependency work, but cautioned against forcing Luna into subagent roles.
OpenAI said GPT-5.6 Luna is 80% cheaper and Terra is 20% cheaper, with lower usage burn in Codex and ChatGPT Work. Sol Fast adds up to 2.5x speed at 2x price, and gateways reflected the new pricing.
Moonshot published Kimi K3 weights, a technical report, and a blog for a 2.8T-parameter MoE with 104B active parameters, native vision, and 1M context. The license adds separate terms for large model-as-a-service providers.
Anthropic released Claude Opus 5 with Fast Mode on paid plans and the API at the Opus 4.8 price. Benchmarks from ARC Prize, Artificial Analysis, Vals AI, and tool vendors put it near or ahead of Fable 5 on several agent and coding tests.
Posts quoting Liang Wenfeng said DeepSeek is targeting low positive API margins while constrained by GPU supply. They also said early-June capacity was about 20,000 H100-equivalent units and that Huawei capacity remains below Nvidia.
Posts citing Anthropic say Fable 5 will stay in Claude Max and Team Premium at 50% of limits. Pro and Team Standard move to credit access with a one-time $100 credit after July 20.
Starting July 20, Fable 5 is included in Claude Max and Team Premium at 50% of plan limits. Pro and Team Standard shift to credit-based access with a one-time $100 credit.
Anthropic said it resolved an issue that made Fable 5 unavailable in claude.ai and Claude Code after users reported lost access or credit prompts. ClaudeDevs said affected extra-usage customers would receive refunds plus matching credits.
OpenAI said Codex and ChatGPT Work reached 9M active users and received another limit reset while reliability work continued. Users still reported weekly caps after long GPT-5.6 Sol coding runs.
Cursor users reported unintended Claude Opus, Sonnet, Fable, or API calls after selecting other settings. Reports included 22.6M-token burns, hard usage stops, and unexpected bills.
OpenAI said 7M active users now use Codex and ChatGPT Work, and paid users reported an extra reset card in web and mobile settings. Other reports said GPT-5.6 Sol context changes may reduce Codex burn.