Skip to content
AI Primer
TOPIC50 stories

Cost Optimization

Reducing inference spend and improving unit economics.

NEWS10th August
Composio and Ante benchmark coding agent harnesses with 47%–67% success range

Composio and Ante tests reported that the same models behaved very differently by harness. DeepSeek V4 Flash ranged from 47% to 67% task success and $0.019 to $0.104 per task across harnesses.

RELEASE10th August
OpenRouter updates Auto Router with 30 task types and 7-day spend-based routing

OpenRouter upgraded Auto Router to classify prompts into about 30 task types, then route by anonymized 7-day spend share and cost tier. OpenRouter says the max tier beat the old router across five benchmark domains.

NEWS9th August
Microsoft Copilot traces report 87% of LLM calls came from agents

A Microsoft Copilot trace analysis said 87% of LLM calls came from the agent, not direct user turns. Related posts warned token use and web requests can scale far faster than human prompt counts.

NEWS9th August
OpenCode users reportedly average $1.14/day on DeepSeek V4 Flash

OpenCode’s thdxr said Go users spent $1.14 per day on DeepSeek V4 Flash last week. Wafer added a fast OpenRouter route, while Nous extended a 90% discount for the 0731 model.

NEWS8th August
DeepSeek V4 Flash benchmarks claim lower DeepSWE cost than GPT-5.6 Luna

Together says two V4 Flash attempts solved more DeepSWE tasks than one GPT-5.6 Luna attempt for roughly one-third the cost. Practitioners report Flash-0731 results vary sharply by harness and pass count.

NEWS8th August
Vercel details spend caps and anomaly alerts for runaway agent bills

Guillermo Rauch listed shipped safeguards including spend caps, anomaly alerts, recursion protection, billing APIs and DDoS mitigation. The post followed reports of agents looping until queues timed out.

NEWS7th August
DeepSeek V4 Flash benchmarks at 61.4% on ARC-AGI-2 for $0.04 per task

ARC Prize verified DeepSeek V4 Flash at 61.4% on ARC-AGI-2 for $0.04 per task. Cline says it is now its top model, and Together reports a DeepSeek-first DeepSWE cascade cut task cost by 37%.

NEWS7th August
Databricks reports coding-agent token spend is rising exponentially

Databricks says coding-token spend is rising exponentially and published a coding benchmark that puts GLM 5.2, Claude Opus 4.8, and GPT-5.6 Sol on the quality-per-dollar frontier. Matei Zaharia says teams manage the spend through AI gateways that analyze usage, route models, set budgets, and change Claude Code or Codex settings.

RELEASE1w ago
DeepSeek V4 Flash adds Baseten and Together AI serving with 1M-token context

Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.

WORKFLOW1w ago
Codex users route tasks across GPT-5.6 Sol, Terra, and Luna to cut token cost

Practitioners reported better Codex multi-agent runs by raising concurrency and splitting work across Sol, Terra, and Luna. One workflow sends deploy tasks to Luna Max to preserve Sol tokens.

NEWS1w ago
Cline raises free DeepSeek Flash quota 3x for coding agents

Cline tripled its free quota while Nous discounted Flash 0731 by 90% for a week and OpenHands offered free cloud use. OpenCode reported 8T Flash tokens on Aug. 1.

NEWS1w ago
DeepSeek V4 Flash benchmarks show cheaper tokens but 3x SWE-Bench task cost

New tests showed DeepSeek V4 Flash as cheaper per token and faster on some serving paths. Ramp said it cost 3x more than GPT-5.6 Luna per SWE-Bench task because it used more turns.

WORKFLOW1w ago
Sol-advisor routes Codex tasks across GPT-5.6 Sol, Luna, and Terra

Sol-advisor routes Codex tasks through GPT-5.6 Sol, Luna, and Terra, while users compare Luna Max as a lower-cost reasoning setting. Early reports say small routing tests need larger benchmarks.

NEWS1w ago
OpenAI cuts GPT-5.6 Luna pricing by 80%

OpenAI said GPT-5.6 Luna pricing fell 80%, while Terra fell 20%. Codex users recommended max reasoning for linting, tests, and dependency work, but cautioned against forcing Luna into subagent roles.

NEWS1w ago
OpenAI cuts GPT-5.6 Luna API prices by 80%

OpenAI said GPT-5.6 Luna is 80% cheaper and Terra is 20% cheaper, with lower usage burn in Codex and ChatGPT Work. Sol Fast adds up to 2.5x speed at 2x price, and gateways reflected the new pricing.

NEWS2w ago
OpenAI says GPT-5.6 Sol cuts model-serving costs by 20%

OpenAI says it used GPT-5.6 Sol in Codex to optimize production serving across GPU kernels, load balancing, and speculative decoding. The company reports a 20% end-to-end cost reduction.

NEWS2w ago
Composio benchmarks Kimi K3 harnesses with $0.22–$2 per-task cost swing

Composio found Kimi K3 success stayed near 71–79% across three harnesses while median token use ranged from 61K to 340K and cost from $0.22 to $2 per task. Cline also shipped a Kimi K3 CLI update.

RELEASE2w ago
Gemini API adds token budget caps for Managed Agents

Google added token budget caps and other controls for Managed Agents in the Gemini API. The release also adds sandbox hooks, cron triggers, model configuration, free-tier support, and Gemini 3.6 Flash defaults.

WORKFLOW2w ago
Developers report under-10% unattended completion for coding agents

Robert C. Martin argued tangled code confuses agents and promoted agent-led refactoring with tests. Other reports cited high Sol spend, under-10% unattended completion, and failures once bugs or complexity appear.

NEWS3w ago
Liang Wenfeng reportedly frames DeepSeek roadmap around scarce GPU supply

Posts quoting Liang Wenfeng said DeepSeek is targeting low positive API margins while constrained by GPU supply. They also said early-June capacity was about 20,000 H100-equivalent units and that Huawei capacity remains below Nvidia.

RELEASE3w ago
Cursor Router adds Cost mode for Teams and Enterprise coding requests

Cursor launched Router for Teams and Enterprise with Intelligence, Balance, and Cost modes. Cursor says the router can select models per request and cut costs by 60%, with admin controls for businesses.

RELEASE3w ago
Google ships Gemini 3.6 Flash and 3.5 Flash-Lite to serving platforms

Gemini 3.6 Flash and 3.5 Flash-Lite went live on OpenRouter, Venice, Hyperbrowser, and Google surfaces. Early benchmarks show lower token use and cost, but mixed document and coding results.

RELEASE3w ago
Martian launches Ship beta with 50% lower-cost inference target

Martian launched Ship, a beta endpoint that sits between an app and its reference model and routes each request through cheaper paths while aiming to preserve behavior. Martian says the beta targets 50% lower cost.

NEWS3w ago
METR introduces expenditure horizon for cost-aware agent evals

METR introduced expenditure horizon, a method that compares human and agent performance as a function of spend on continuously scored tasks. A correction estimates human labor returns near $2.5K per 1% optimization, making cost curves central to the evaluation.

NEWS3w ago
Artificial Analysis reports Kimi K3 averages 56.4 minutes on AA-Briefcase

Artificial Analysis reports Kimi K3 averages 56.4 minutes, 83 turns, and 120k output tokens per AA-Briefcase task. Kilo also found UI-build outputs close to Claude Fable 5 at 29% of the cost.

WORKFLOW3w ago
Engineers replace broad agent loops with scoped workflows and SWE-bench harnesses

Practitioner threads favored narrow tasks, deterministic steps, task-specific harnesses and explicit acceptance criteria over open-ended agent loops. One Codex workflow reported 58.6%-87.0% token cuts on SWE-bench task sets.

NEWS3w ago
Kimi K3 benchmarks last at 53/67 in AlphaSignal repair harness

AlphaSignal's repair harness put Kimi K3 last at 53/67, while other tests ranked it high on DeepSWE, Vibe Code Bench, legal work, cyber, and CUDA kernels. Cost often beat Fable, but speed lagged.

NEWS3w ago
Kimi K3 ranks #5 on Artificial Analysis as engineers dispute coding cost

Kimi K3 posted strong coding results, including rank #5 on Artificial Analysis and #3 on DeepSWE. Engineers disputed whether its lower token price offsets higher token use and slower throughput.

NEWS4w ago
Cursor users report agents switching to costly Claude Opus, Sonnet, Fable, or API calls

Cursor users reported unintended Claude Opus, Sonnet, Fable, or API calls after selecting other settings. Reports included 22.6M-token burns, hard usage stops, and unexpected bills.

RELEASE4w ago
Devin Fusion adds Fable 5 to cut coding-agent task costs

Cognition said Devin Fusion now uses Fable 5 and saw lower cost per task than Opus 4.8. Practitioners cited Fable-led delegation patterns that cut token use, with caveats on serial debugging.

NEWS4w ago
Coding Agent Index ranks cheaper configs near the top

Fresh runs and charts put GPT-5.6 Sol high on SWE-Bench Pro and Design Arena, while Coding Agent Index and Amp reports emphasized cheaper strong configs. Results vary by harness, effort tier, and agent setup.

NEWS4w ago
Devin exposes Fusion router with claimed 35% lower frontier-model cost

Devin exposed routers such as Fusion, claiming frontier performance at 35% lower cost, while Databricks argued for smart routing in agent harnesses. New charts put Grok 4.5 and Muse Spark 1.1 near the coding-cost frontier.

NEWS4w ago
GPT-5.6 Sol ranks near top of DeepSWE and coding evals at lower reported cost

New benchmark posts put GPT-5.6 Sol at or near the top of DeepSWE and several coding/context evals. Cost reports placed Luna on the efficiency frontier, while Amp said replacing Opus with GPT-5.6 cut its average model costs ~50%.

NEWS4w ago
Claude Code users report quota burn and usage-meter failures

Reddit users reported Claude Code subagents hanging while burning quota, inconsistent usage meters, oversized contexts, slow desktop output, and a verify skill consuming a full limit. The common issue was unreliable usage accounting.

NEWS4w ago
Muse Spark 1.1 benchmarks near frontier agent models at $1.25/M input

Meta and third-party benchmark posts put Muse Spark 1.1 near frontier coding and agent models at $1.25/M input and $4.25/M output. Results included Vals AI agent tasks, Code Arena Frontend #9, and an AA Coding Agent Index score of 69.

NEWS4w ago
Early benchmarks rank GPT-5.6 Sol near Fable 5 at lower cost

ARC Prize, Artificial Analysis, CursorBench, and other tests reported strong GPT-5.6 Sol results, especially in coding-agent tasks. Results were uneven, with smaller gains in document parsing and some UI or puzzle evals.

WORKFLOW1mo ago
Bun Rust rewrite becomes AI-assisted coding case study with estimated $165k API cost

Simon Willison highlighted Jarred Sumner’s Bun rewrite from Zig to Rust as an AI-assisted workflow using trial runs, dynamic planning, adversarial review, and verification. One estimate put API token cost near $165k.

WORKFLOW1mo ago
ClaudeDevs reports Sonnet 5 + Fable 5 advisor hits ~92% SWE-bench Pro score at ~63% price

ClaudeDevs reports Sonnet 5 with a Fable 5 advisor reached ~92% of Fable 5's SWE-bench Pro score at ~63% of the price. Other builders route implementation to Sonnet, Codex, GPT-5.5, or GLM workers.

NEWS1mo ago
OpenRouter benchmarks 1,730 visual-reasoning questions on low-detail image costs

OpenRouter tested 1,730 visual-reasoning questions across five models and found low-detail images often reduced accuracy while increasing reasoning-token spend. Caps on reasoning effort had the biggest billing impact.

WORKFLOW1mo ago
OpenRouter claims 24x inference-cost savings with MCP model routing

OpenRouter published an MCP workflow that it says cut inference costs 24x at comparable quality. The MCP lets the model choose providers using codebase context plus OpenRouter benchmark, aggregate-usage, and live-performance data.

NEWS1mo ago
Fable users report $130 prompts and quota drain

Fable users reported cost escalations from $300/day estimates to a single high-effort prompt above $130 and quota-drain complaints on Reddit. Users also reported automatic Opus fallback and safety refusals tied to bio/cyber safeguards.

WORKFLOW1mo ago
Claude Code user estimates subagent prompt caching raised spend by 8%

One Claude Code user parsed 95 sessions and estimated subagent prompt caching made total spend about 8% high. pxpipe separately rendered dense text as images to cut context cost, with exactness tradeoffs.

RELEASE1mo ago
Condense.chat opens Adeline 1 proxy for 9% agent-loop compaction

Condense.chat opened a compression proxy that strips tokens with Helene 1 and compacts settled agent loops with Adeline 1 to about 9% of their size. The service claims 100M saved tokens and 3× plan extension for Claude or Codex users, so test it on non-sensitive workflows first.

NEWS1mo ago
Ramp introduces PorTAL with half-cost LoRA porting across Qwen and Gemma models

Ramp published PorTAL, a method that learns a reusable task representation once and recalibrates only a thin converter when moving that task to a new base model. In reported Qwen and Gemma experiments, it matched per-task LoRA accuracy while cutting data and cost roughly in half.

NEWS1mo ago
The Information reports OpenAI cuts inference costs by more than 50% on some models

Multiple summaries of The Information report said OpenAI found inference optimizations that more than halved costs on some existing models. If that holds, it changes the margin, pricing, and usage-limit math behind ChatGPT and API serving even before new model releases arrive.

RELEASE1mo ago
Hermes Agent updates web extraction with 60x faster reads and 49x lower cost

Nous updated Hermes Agent web extraction to skip the old summarizer loop, pass cleaner content directly to the model, and page large documents on demand. The change is claimed to cut read latency by up to 60x and cost by 49x, so teams should compare output quality before adopting it.

RELEASE1mo ago
Junior adds memory and cuts one analytics task from 3m to 1m

Junior’s first memory system cut one analytics task from about 3 minutes to 1 minute in early tests, with tokens down two-thirds and tool calls down 60%. The feature moves persistent task learning into the agent loop, though the results are still internal.

RELEASE1mo ago
Kilo Code launches Auto Efficient routing with KiloBench model selection

Kilo Code added an Auto Efficient mode that routes each request to the cheapest model that clears its benchmark bar using public KiloBench results. The router stays session-aware and falls back to stronger paid models when confidence is low.

NEWS1mo ago
GLM-5.2 adds Perplexity Agent API and Droid support on Baseten at >280 TPS

GLM-5.2 added Perplexity Agent API, Droid, and more hosting options, while Baseten reported over 280 TPS and sub-0.8s TTFT. Builders should watch the cost and benchmark data as it moves into production agent stacks.

NEWS1mo ago
Wafer claims GLM-5.2 hits 222 tok/s and 12.6s end-to-end

Wafer said its GLM-5.2 deployment leads Artificial Analysis on throughput and latency, and priced usage at $1.20 input and $4.10 output per million tokens. Compare serverless and dedicated endpoints if you need speed at scale.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.