Skip to content
AI Primer
TOPIC50 stories

Pricing & limits

Pricing changes, quotas, rate limits, and plan/tier changes for AI dev tools and APIs.

NEWS17th August
OpenCode Go adds $30 usage credit to its $10 plan

OpenCode revised its Go plan to include $30 in usage for a $10 subscription. The change follows a provider price increase and the company’s effort to secure lower-cost capacity.

NEWS16th August
OpenCode Go revises limits after DeepSeek price increase

OpenCode says it revised Go limits after DeepSeek raised prices. Its operator is testing hosting configurations intended to bring DeepSeek service closer to its prior price point.

RELEASE16th August
Codex opens 1M-token GPT-5.6 Sol context for ChatGPT subscribers

Codex now lets ChatGPT subscribers enable a 1 million-token context window for GPT-5.6 Sol. Automatic compaction starts at 900,000 tokens, and tokens beyond the default window count double against limits.

NEWS1w ago
OpenCode users reportedly average $1.14/day on DeepSeek V4 Flash

OpenCode’s thdxr said Go users spent $1.14 per day on DeepSeek V4 Flash last week. Wafer added a fast OpenRouter route, while Nous extended a 90% discount for the 0731 model.

NEWS1w ago
Vercel details spend caps and anomaly alerts for runaway agent bills

Guillermo Rauch listed shipped safeguards including spend caps, anomaly alerts, recursion protection, billing APIs and DDoS mitigation. The post followed reports of agents looping until queues timed out.

RELEASE1w ago
Seedance 2.5 ships in ComfyUI with 30-second runs and timeline shot control

ComfyUI says Seedance 2.5 is live via Partner Nodes with 30-second runs, up to 50 references, timeline shot control, editing, and multilingual lip sync. Fal, Pika, Venice, and other tools also added access.

WORKFLOW2w ago
Cursor users report hard-to-audit coding-agent runs and hidden routing

Cursor users say AI IDE agent runs are hard to audit and hard to constrain. Reddit threads cite hidden model routing, cache charges, destructive SQL migrations, rules folders, runbooks, and context-trimming pipelines.

NEWS2w ago
Cline raises free DeepSeek Flash quota 3x for coding agents

Cline tripled its free quota while Nous discounted Flash 0731 by 90% for a week and OpenHands offered free cloud use. OpenCode reported 8T Flash tokens on Aug. 1.

NEWS2w ago
OpenAI cuts GPT-5.6 Luna pricing by 80%

OpenAI said GPT-5.6 Luna pricing fell 80%, while Terra fell 20%. Codex users recommended max reasoning for linting, tests, and dependency work, but cautioned against forcing Luna into subagent roles.

NEWS2w ago
Claude Code users report Fable 5 and Opus 5 burn 5-hour limits in 6–10 minutes

Claude Code users reported Fable 5 and Opus 5 sessions exhausting five-hour usage windows in about 6–10 minutes after automated tool calls. The reports tie the failures to agent loops and rate-limit economics, including one Reddit claim of 10.26M tokens, 15 calls, and no edits.

NEWS2w ago
OpenAI cuts GPT-5.6 Luna API prices by 80%

OpenAI said GPT-5.6 Luna is 80% cheaper and Terra is 20% cheaper, with lower usage burn in Codex and ChatGPT Work. Sol Fast adds up to 2.5x speed at 2x price, and gateways reflected the new pricing.

RELEASE3w ago
Moonshot releases Kimi K3 open weights with 2.8T-parameter MoE

Moonshot published Kimi K3 weights, a technical report, and a blog for a 2.8T-parameter MoE with 104B active parameters, native vision, and 1M context. The license adds separate terms for large model-as-a-service providers.

RELEASE3w ago
Anthropic ships Claude Opus 5 to paid plans and API at Opus 4.8 price

Anthropic released Claude Opus 5 with Fast Mode on paid plans and the API at the Opus 4.8 price. Benchmarks from ARC Prize, Artificial Analysis, Vals AI, and tool vendors put it near or ahead of Fable 5 on several agent and coding tests.

NEWS4w ago
Liang Wenfeng reportedly frames DeepSeek roadmap around scarce GPU supply

Posts quoting Liang Wenfeng said DeepSeek is targeting low positive API margins while constrained by GPU supply. They also said early-June capacity was about 20,000 H100-equivalent units and that Huawei capacity remains below Nvidia.

NEWS4w ago
Posts say Anthropic makes Fable 5 permanent on Claude Max and Team Premium at 50% limits

Posts citing Anthropic say Fable 5 will stay in Claude Max and Team Premium at 50% of limits. Pro and Team Standard move to credit access with a one-time $100 credit after July 20.

NEWS4w ago
Anthropic adds Fable 5 to Max and Team Premium at 50% limits

Starting July 20, Fable 5 is included in Claude Max and Team Premium at 50% of plan limits. Pro and Team Standard shift to credit-based access with a one-time $100 credit.

NEWS4w ago
Anthropic fixes Fable 5 selection outage and issues refunds

Anthropic said it resolved an issue that made Fable 5 unavailable in claude.ai and Claude Code after users reported lost access or credit prompts. ClaudeDevs said affected extra-usage customers would receive refunds plus matching credits.

NEWS1mo ago
Cursor users report agents switching to costly Claude Opus, Sonnet, Fable, or API calls

Cursor users reported unintended Claude Opus, Sonnet, Fable, or API calls after selecting other settings. Reports included 22.6M-token burns, hard usage stops, and unexpected bills.

NEWS1mo ago
OpenAI resets Codex and ChatGPT Work limits after 9M active users

OpenAI said Codex and ChatGPT Work reached 9M active users and received another limit reset while reliability work continued. Users still reported weekly caps after long GPT-5.6 Sol coding runs.

NEWS1mo ago
OpenAI adds banked reset card for Codex paid users

OpenAI said 7M active users now use Codex and ChatGPT Work, and paid users reported an extra reset card in web and mobile settings. Other reports said GPT-5.6 Sol context changes may reduce Codex burn.

NEWS1mo ago
OpenAI reverts Codex GPT-5.6 Sol context limit to 272k after overcharging

OpenAI said GPT-5.6 Sol's 372k context in Codex charged more usage than intended, so it reverted Codex to 272k. It also removed a five-hour cap, reset some rates, and passed inference savings into more subscription usage.

NEWS1mo ago
Anthropic extends Fable 5 paid-plan access through July 19

Anthropic extended Fable 5 on paid plans and raised Claude Code weekly limits by 50% through July 19 while keeping the half-week cap. Users still described the rolling access changes as disruptive.

RELEASE1mo ago
Genspark and OpenClaw add Grok 4.5 for coding-agent workflows

Genspark and OpenClaw added Grok 4.5 after xAI's launch, extending the model into more coding-agent workflows. Follow-up evidence covered AA-Briefcase and Terminal-Bench results, a Composio credential-audit run, and SuperGrok usage-meter reports.

RELEASE1mo ago
Grok 4.5 launches for coding agents at $2/M input and $6/M output

SpaceXAI launched Grok 4.5 in Cursor and several agent tools with $2/M input and $6/M output pricing. Early evals place it near frontier coding models, with 51% on AutomationBench-AA.

NEWS1mo ago
Anthropic extends paid Fable 5 access through July 12 with 50% weekly cap

Anthropic says paid Claude plans keep Fable 5 access until July 12, with Fable use capped at 50% of weekly limits. Users still reported exhausted quotas and multiple Max subscriptions as workarounds.

NEWS1mo ago
Fable 5 users report heavy Claude Code credit burn before July 7 cutoff

A Fable maintainer set the Claude subscription cutoff for 11:59:59pm PT on July 7 as users posted high Claude Code credit usage. One estimate claimed about $2,267 of usage across two $200 accounts in six days.

NEWS1mo ago
Fable 5 users report July 7 Claude Max cutoff and compute-return caveat

Fable 5 users said Claude Max subscription access runs through July 7, with some reporting it may return only when compute is available. The deadline changed planning for Fable-heavy coding runs and fallback options.

WORKFLOW1mo ago
Claude Code user estimates subagent prompt caching raised spend by 8%

One Claude Code user parsed 95 sessions and estimated subagent prompt caching made total spend about 8% high. pxpipe separately rendered dense text as images to cut context cost, with exactness tradeoffs.

NEWS1mo ago
Fable users report $130 prompts and quota drain

Fable users reported cost escalations from $300/day estimates to a single high-effort prompt above $130 and quota-drain complaints on Reddit. Users also reported automatic Opus fallback and safety refusals tied to bio/cyber safeguards.

NEWS1mo ago
Fable 5 users report Opus 4.8 fallbacks and $600 Max quota rotations

Fable 5 users reported Opus 4.8 fallbacks, $600 Max-account rotations, slow browser automation, and token-saving subagents. Watch routing opacity, quota burn, and latency before relying on it for long-running agent work.

NEWS1mo ago
Codex app reportedly leaks GPT-5.6 Sol, Terra, and Luna model names

Codex app code now references GPT-5.6 Sol, Terra, and Luna, while posts claim Sol Ultra reaches 91.9% on TerminalBench at lower cost. Treat release timing, limits, and benchmark claims as unofficial until OpenAI publishes details.

RELEASE1mo ago
Condense.chat opens Adeline 1 proxy for 9% agent-loop compaction

Condense.chat opened a compression proxy that strips tokens with Helene 1 and compacts settled agent loops with Adeline 1 to about 9% of their size. The service claims 100M saved tokens and 3× plan extension for Claude or Codex users, so test it on non-sensitive workflows first.

NEWS1mo ago
Fable 5 users report Opus 4.8 fallbacks, refusals, and $321 sessions

Users posted mixed reports after Anthropic brought Fable 5 back: some sessions stayed on Fable, while others routed most work to Opus 4.8 or stalled mid-run. Watch for routing changes and cost spikes, since reports also mention refusals on ordinary tasks and ad hoc multi-model workarounds.

NEWS1mo ago
Claude Sonnet 5 ranks #3 on Vals and hits 183 turns on AA-Briefcase

Vals and Artificial Analysis published independent Sonnet 5 results a day after launch, placing it just behind Opus 4.8 and Fable 5 while using far more turns than Sonnet 4.6. Lower token pricing did not make agentic tasks cheaper, and some finance benchmarks still triggered refusals.

RELEASE1mo ago
Google releases Nano Banana 2 Lite and Gemini Omni Flash

Google shipped Nano Banana 2 Lite for image generation and Gemini Omni Flash for conversational video generation and editing in the Gemini API and AI Studio. The release sets image generation at about 4 seconds and $0.034 per 1K image, while Omni Flash adds multi-turn video edits at $0.10 per second.

RELEASE1mo ago
Cognition launches Devin Fusion with mid-session routing and 35% lower Fable-class cost

Cognition launched Devin Fusion, a hybrid coding harness that reroutes work mid-task and says it cuts Fable-class cost by 35%. Use it when upfront routing misses late complexity; the router can re-evaluate after investigation starts.

NEWS1mo ago
Codex fixes usage overcounting with one extra banked reset and auto-review rollback

A day after Codex reset limits for weekend drain reports, OpenAI said auto-review, duplicate background suggestions, and retry behavior were compounding usage and issued another full reset. Users also get one extra reset credit within 24 hours while reporting and scheduling fixes roll out.

NEWS1mo ago
OpenRouter reports four open-weight models handle agents; Chinese models hit 45% of traffic

OpenRouter said four open-weight models now handle real agentic workloads, and a JPMorgan report put Chinese models at about 45% of platform traffic. The shift matters because teams are optimizing for price, hosting, and task fit instead of defaulting to frontier APIs.

RELEASE1mo ago
Datalab ranks 95.9% on a 225-document extraction benchmark at under half Reducto's price

Datalab’s balanced extraction mode scored 95.9% on a 225-document benchmark and beat Reducto Deep Extract’s 95.1%, according to Vik Paruchuri. The update also adds citations and reasoning, but the benchmark and price comparison are vendor-reported.

NEWS1mo ago
Codex fixes quota drain tied to fraud overflagging with an account-wide usage reset

OpenAI said Codex accounts were seeing faster usage draining than intended because abuse and fraud checks were overflagging some sessions, then issued a usage reset for all users. It matters because paid Codex workflows were losing quota unexpectedly mid-run, directly affecting reliability and cost.

RELEASE1mo ago
Seedance 2.0 Mini launches on Venice, ComfyUI, and Pika MCP with 15s 720p video

A day after Seedance 2.0's 4K rollout story, partners began shipping the cheaper Seedance 2.0 Mini across Venice, ComfyUI, and Pika MCP. The 15-second 720p variant with native audio gives video workflows a lower-cost path than the flagship model.

NEWS1mo ago
Claude Tag users report token billing and shared-memory concerns

A day after Claude Tag launched, engineers raised token billing, lock-in, and shared-memory concerns while Anthropic described its agent-identity model. Watch how Claude behaves in shared Slack channels, where it uses its own credentials and scoped access.

RELEASE1mo ago
Vercel AI Gateway adds GLM-5.2 Fast at 150-250 tok/s

Vercel and Wafer launched a serverless GLM-5.2 endpoint on AI Gateway with 1M context and published pricing. Teams get a high-throughput open-model option inside an existing gateway instead of managing GLM inference directly.

RELEASE1mo ago
Kilo Code launches Auto Efficient routing with KiloBench model selection

Kilo Code added an Auto Efficient mode that routes each request to the cheapest model that clears its benchmark bar using public KiloBench results. The router stays session-aware and falls back to stronger paid models when confidence is low.

NEWS2mo ago
GLM-5.2 ranks #1 on DeepSWE with 44% pass@1

Independent results put GLM-5.2 at the top of the open-model DeepSWE board and near the top on debate and post-train evals. Watch token use and long reasoning traces, which can offset its headline price advantage.

NEWS2mo ago
Engineers compare GLM-5.2 local builds: $10k Mac Studio, 17 tok/s, and 2-bit quant tradeoffs

Practitioners published concrete GLM-5.2 self-host numbers, from Mac Studio and 4090-class setups to annualized power and hardware costs. That matters because open weights now offer privacy and rate-limit control, but quant quality, electricity, and latency still keep hosted APIs cheaper for many teams.

NEWS2mo ago
Engineers report GLM-5.2 matches near-Opus planning at about 1/10 the price

Independent tests put GLM-5.2 near Opus 4.8 and GPT-5.5 on planning and coding, and users shared Claude Code, BrowserCode, dcode, and local-serving recipes. It matters because many engineers are treating it as a daily-driver option for text-heavy coding, though teams still report weaker vision and provider limits.

RELEASE2mo ago
Kilo Code adds Terminal Bench scores and average attempt cost to model picker

Kilo Code now shows Terminal Bench completion rate and average attempt cost directly in model details inside its CLI and VS Code extension. It matters because the numbers come from Kilo's own harness and retry logic rather than public leaderboard scaffolds.

RELEASE2mo ago
Moonshot releases Kimi K2.7 Code HighSpeed at 180 tok/s with 2x API pricing

Moonshot rolled out HighSpeed for Kimi K2.7 Code, claiming about 180 tok/s on coding tasks, up to 260 tok/s on shorter contexts, and roughly 6x speedups. Watch the tight capacity limits and mixed benchmark results, and budget for the 2x pricing if you want the faster mode.

NEWS2mo ago
Anthropic delays Claude Agent SDK credit shift for claude -p and third-party apps

Anthropic paused a same-day policy change that would have moved Claude Agent SDK, claude -p, and third-party SDK apps onto separate monthly credits. Existing subscription-backed workflows continue unchanged for now, but teams should watch for the redesigned billing plan.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.