Cost Optimization
Reducing inference spend and improving unit economics.
Stories
Filter storiesExperiments reported by Maxime Rivest found task-specific small classifiers outperforming frontier models on specialized decisions. One result says a 4-million-parameter BERT Tiny model beat Opus and Kimi after training on 10,000 examples.
An independent DeepSWE test found Jev Router roughly matched GPT-6 Astra low on results but cost slightly more and took nearly five times as long. Practitioners also argue that request-level routing misses repository context and cache costs.
OpenRouter launched Jev Router, which selects a model and reasoning effort per turn while weighing the cost of losing cached context. OpenRouter reports 237 of 423 tasks solved; invalid Jev outputs or timeouts fail without fallback.
OpenAI released GPT-6 Sol and GPT-6 Luna at API prices roughly half those of their GPT-5.6 predecessors. The models add controllable prompt-cache breakpoints and are rolling out in Codex and ChatGPT Work.
Teknium’s public evaluation says a Jev compaction strategy removes tool calls and eventually stops yielding savings. Repeated compaction can invalidate caches and increase total token costs, according to the critique.
Stagehand says adding Jev to its browser primitives cut median Act latency from 1.97 seconds to 0.46 seconds. Jev handles page-level choices and falls back to an LLM when uncertain.
Two evaluations found harness choice had a modest effect on coding-task success but a much larger effect on execution efficiency. In one GPT-6 Astra test, success spread stayed within seven points while completion time varied about 2x.
MiMo is reportedly livestreaming RL training for its V2.6 Pro and Flash models, publishing batch data, harness composition, reward curves, and infrastructure metrics. Reported cost figures list the trillion-parameter Pro run at about $493,000.
Polylane says it replaced role-specific sub-agents with one main agent and improved quality while reducing latency and cost. The report is a practitioner case study, not a general benchmark.
Fusion keeps a lead model in control while routing execution to a cheaper model in Devin CLI. Cognition reported 39% lower benchmark cost, and Artificial Analysis measured 43% lower cost with 31% faster runs at near-frontier scores.
Cognition says its SWE-2 coding model scored 50.0% on FrontierCode 1.1 Main, matching Fable 5.1 at 64% lower cost. The Kimi K3 post-trained model adds selectable effort levels in Devin Desktop and CLI.
Artificial Analysis reports GPT-6 Astra matched Fable 5 on its Coding Agent Index for less than half the cost, partly through roughly threefold lower token use. Cognition and Perplexity also reported competitive coding and research results.
Cohere released Parse 5, a document parser that returns machine-readable text, tables, forms, images, and bounding boxes. Cohere reports a 79.2 ParseBench score and prices it at $1.50 per 1,000 pages.
Glean says its runtime routes enterprise-agent work across more than 40 models using company context. It reports $0.58 per query and 78% user preference over Claude Cowork in a 180-person benchmark.
Together reports GLM-5.3 solved 87.6% of DeepSWE after four attempts for about $16, versus Fable 5 at 69.7% for $21.63. It estimates equal $100 budgets yield roughly 17 solved tasks for GLM-5.3 and three for Fable 5.
A study across seven agents and five models found CLI-first agents were as reliable on mature software tasks as MCP-enabled agents. The CLI setups cost 5–28 times less in the reported experiments.
OpenAI cut GPT-5.6 Sol API input and output rates from $5 and $30 to $4 and $20 per million tokens for three months. Subscription usage remains unchanged.
ARC Prize verified Gemini 3.7 Flash at 84.6% on ARC-AGI-2, at a reported $0.25 per task. Artificial Analysis’ AnalystAgent benchmark placed it at 60%, ahead of Claude Opus 5 and GPT-5.5.
OpenCode says it revised Go limits after DeepSeek raised prices. Its operator is testing hosting configurations intended to bring DeepSeek service closer to its prior price point.
Composio and Ante tests reported that the same models behaved very differently by harness. DeepSeek V4 Flash ranged from 47% to 67% task success and $0.019 to $0.104 per task across harnesses.
OpenRouter upgraded Auto Router to classify prompts into about 30 task types, then route by anonymized 7-day spend share and cost tier. OpenRouter says the max tier beat the old router across five benchmark domains.
A Microsoft Copilot trace analysis said 87% of LLM calls came from the agent, not direct user turns. Related posts warned token use and web requests can scale far faster than human prompt counts.
OpenCode’s thdxr said Go users spent $1.14 per day on DeepSeek V4 Flash last week. Wafer added a fast OpenRouter route, while Nous extended a 90% discount for the 0731 model.
Together says two V4 Flash attempts solved more DeepSWE tasks than one GPT-5.6 Luna attempt for roughly one-third the cost. Practitioners report Flash-0731 results vary sharply by harness and pass count.
Guillermo Rauch listed shipped safeguards including spend caps, anomaly alerts, recursion protection, billing APIs and DDoS mitigation. The post followed reports of agents looping until queues timed out.
ARC Prize verified DeepSeek V4 Flash at 61.4% on ARC-AGI-2 for $0.04 per task. Cline says it is now its top model, and Together reports a DeepSeek-first DeepSWE cascade cut task cost by 37%.
Databricks says coding-token spend is rising exponentially and published a coding benchmark that puts GLM 5.2, Claude Opus 4.8, and GPT-5.6 Sol on the quality-per-dollar frontier. Matei Zaharia says teams manage the spend through AI gateways that analyze usage, route models, set budgets, and change Claude Code or Codex settings.
Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.
Cline tripled its free quota while Nous discounted Flash 0731 by 90% for a week and OpenHands offered free cloud use. OpenCode reported 8T Flash tokens on Aug. 1.
Practitioners reported better Codex multi-agent runs by raising concurrency and splitting work across Sol, Terra, and Luna. One workflow sends deploy tasks to Luna Max to preserve Sol tokens.
New tests showed DeepSeek V4 Flash as cheaper per token and faster on some serving paths. Ramp said it cost 3x more than GPT-5.6 Luna per SWE-Bench task because it used more turns.
Sol-advisor routes Codex tasks through GPT-5.6 Sol, Luna, and Terra, while users compare Luna Max as a lower-cost reasoning setting. Early reports say small routing tests need larger benchmarks.
OpenAI said GPT-5.6 Luna pricing fell 80%, while Terra fell 20%. Codex users recommended max reasoning for linting, tests, and dependency work, but cautioned against forcing Luna into subagent roles.
OpenAI said GPT-5.6 Luna is 80% cheaper and Terra is 20% cheaper, with lower usage burn in Codex and ChatGPT Work. Sol Fast adds up to 2.5x speed at 2x price, and gateways reflected the new pricing.
OpenAI says it used GPT-5.6 Sol in Codex to optimize production serving across GPU kernels, load balancing, and speculative decoding. The company reports a 20% end-to-end cost reduction.
Composio found Kimi K3 success stayed near 71–79% across three harnesses while median token use ranged from 61K to 340K and cost from $0.22 to $2 per task. Cline also shipped a Kimi K3 CLI update.
Google added token budget caps and other controls for Managed Agents in the Gemini API. The release also adds sandbox hooks, cron triggers, model configuration, free-tier support, and Gemini 3.6 Flash defaults.
Robert C. Martin argued tangled code confuses agents and promoted agent-led refactoring with tests. Other reports cited high Sol spend, under-10% unattended completion, and failures once bugs or complexity appear.
Posts quoting Liang Wenfeng said DeepSeek is targeting low positive API margins while constrained by GPU supply. They also said early-June capacity was about 20,000 H100-equivalent units and that Huawei capacity remains below Nvidia.
Cursor launched Router for Teams and Enterprise with Intelligence, Balance, and Cost modes. Cursor says the router can select models per request and cut costs by 60%, with admin controls for businesses.
Gemini 3.6 Flash and 3.5 Flash-Lite went live on OpenRouter, Venice, Hyperbrowser, and Google surfaces. Early benchmarks show lower token use and cost, but mixed document and coding results.
Martian launched Ship, a beta endpoint that sits between an app and its reference model and routes each request through cheaper paths while aiming to preserve behavior. Martian says the beta targets 50% lower cost.
METR introduced expenditure horizon, a method that compares human and agent performance as a function of spend on continuously scored tasks. A correction estimates human labor returns near $2.5K per 1% optimization, making cost curves central to the evaluation.
Artificial Analysis reports Kimi K3 averages 56.4 minutes, 83 turns, and 120k output tokens per AA-Briefcase task. Kilo also found UI-build outputs close to Claude Fable 5 at 29% of the cost.
Practitioner threads favored narrow tasks, deterministic steps, task-specific harnesses and explicit acceptance criteria over open-ended agent loops. One Codex workflow reported 58.6%-87.0% token cuts on SWE-bench task sets.
AlphaSignal's repair harness put Kimi K3 last at 53/67, while other tests ranked it high on DeepSWE, Vibe Code Bench, legal work, cyber, and CUDA kernels. Cost often beat Fable, but speed lagged.
Kimi K3 posted strong coding results, including rank #5 on Artificial Analysis and #3 on DeepSWE. Engineers disputed whether its lower token price offsets higher token use and slower throughput.
Cursor users reported unintended Claude Opus, Sonnet, Fable, or API calls after selecting other settings. Reports included 22.6M-token burns, hard usage stops, and unexpected bills.
Cognition said Devin Fusion now uses Fable 5 and saw lower cost per task than Opus 4.8. Practitioners cited Fable-led delegation patterns that cut token use, with caveats on serial debugging.
Fresh runs and charts put GPT-5.6 Sol high on SWE-Bench Pro and Design Arena, while Coding Agent Index and Amp reports emphasized cheaper strong configs. Results vary by harness, effort tier, and agent setup.