Skip to content
AI Primer
TOPIC50 stories

Cost Optimization

Reducing inference spend and improving unit economics.

WORKFLOW26th September
4M-parameter BERT Tiny reportedly beats Opus and Kimi after training on 10,000 examples

Experiments reported by Maxime Rivest found task-specific small classifiers outperforming frontier models on specialized decisions. One result says a 4-million-parameter BERT Tiny model beat Opus and Kimi after training on 10,000 examples.

NEWS26th September
Independent DeepSWE test reports Jev Router nearly 5x slower than GPT-6 Astra low

An independent DeepSWE test found Jev Router roughly matched GPT-6 Astra low on results but cost slightly more and took nearly five times as long. Practitioners also argue that request-level routing misses repository context and cache costs.

RELEASE25th September
OpenRouter launches Jev Router for cache-aware model routing

OpenRouter launched Jev Router, which selects a model and reasoning effort per turn while weighing the cost of losing cached context. OpenRouter reports 237 of 423 tasks solved; invalid Jev outputs or timeouts fail without fallback.

RELEASE22nd September
OpenAI releases GPT-6 Sol and GPT-6 Luna at roughly half GPT-5.6 API prices

OpenAI released GPT-6 Sol and GPT-6 Luna at API prices roughly half those of their GPT-5.6 predecessors. The models add controllable prompt-cache breakpoints and are rolling out in Codex and ChatGPT Work.

WORKFLOW1w ago
Teknium reports Jev compaction can increase token costs

Teknium’s public evaluation says a Jev compaction strategy removes tool calls and eventually stops yielding savings. Repeated compaction can invalidate caches and increase total token costs, according to the critique.

WORKFLOW1w ago
Jev cuts Stagehand's median Act latency from 1.97s to 0.46s, Stagehand says

Stagehand says adding Jev to its browser primitives cut median Act latency from 1.97 seconds to 0.46 seconds. Jev handles page-level choices and falls back to an LLM when uncertain.

NEWS1w ago
Coding-agent harness benchmarks report about 2x completion-time spread on GPT-6 Astra

Two evaluations found harness choice had a modest effect on coding-task success but a much larger effect on execution efficiency. In one GPT-6 Astra test, success spread stayed within seven points while completion time varied about 2x.

NEWS1w ago
MiMo reportedly streams V2.6 Pro and Flash RL training metrics

MiMo is reportedly livestreaming RL training for its V2.6 Pro and Flash models, publishing batch data, harness composition, reward curves, and infrastructure metrics. Reported cost figures list the trillion-parameter Pro run at about $493,000.

WORKFLOW1w ago
Polylane says one agent improved quality while cutting latency and cost

Polylane says it replaced role-specific sub-agents with one main agent and improved quality while reducing latency and cost. The report is a practitioner case study, not a general benchmark.

RELEASE2w ago
Cognition launched Devin Fusion to cut coding-agent costs

Fusion keeps a lead model in control while routing execution to a cheaper model in Devin CLI. Cognition reported 39% lower benchmark cost, and Artificial Analysis measured 43% lower cost with 31% faster runs at near-frontier scores.

RELEASE2w ago
Cognition releases SWE-2 with lower FrontierCode costs

Cognition says its SWE-2 coding model scored 50.0% on FrontierCode 1.1 Main, matching Fable 5.1 at 64% lower cost. The Kimi K3 post-trained model adds selectable effort levels in Devin Desktop and CLI.

NEWS3w ago
Artificial Analysis reports GPT-6 Astra matches Fable 5 coding at under half the cost

Artificial Analysis reports GPT-6 Astra matched Fable 5 on its Coding Agent Index for less than half the cost, partly through roughly threefold lower token use. Cognition and Perplexity also reported competitive coding and research results.

RELEASE4w ago
Cohere releases Parse 5 with 79.2 ParseBench score

Cohere released Parse 5, a document parser that returns machine-readable text, tables, forms, images, and bounding boxes. Cohere reports a 79.2 ParseBench score and prices it at $1.50 per 1,000 pages.

NEWS4w ago
Glean says runtime routing cuts enterprise-agent token costs by 81%

Glean says its runtime routes enterprise-agent work across more than 40 models using company context. It reports $0.58 per query and 78% user preference over Claude Cowork in a 180-person benchmark.

NEWS1mo ago
Together reports GLM-5.3 solves 87.6% of DeepSWE work for about $16

Together reports GLM-5.3 solved 87.6% of DeepSWE after four attempts for about $16, versus Fable 5 at 69.7% for $21.63. It estimates equal $100 budgets yield roughly 17 solved tasks for GLM-5.3 and three for Fable 5.

NEWS1mo ago
Study finds CLI-first agents cost 5–28x less than MCP agents

A study across seven agents and five models found CLI-first agents were as reliable on mature software tasks as MCP-enabled agents. The CLI setups cost 5–28 times less in the reported experiments.

NEWS1mo ago
OpenAI cuts GPT-5.6 Sol API rates to $4 input and $20 output

OpenAI cut GPT-5.6 Sol API input and output rates from $5 and $30 to $4 and $20 per million tokens for three months. Subscription usage remains unchanged.

NEWS1mo ago
ARC Prize verifies Gemini 3.7 Flash at 84.6% on ARC-AGI-2

ARC Prize verified Gemini 3.7 Flash at 84.6% on ARC-AGI-2, at a reported $0.25 per task. Artificial Analysis’ AnalystAgent benchmark placed it at 60%, ahead of Claude Opus 5 and GPT-5.5.

NEWS1mo ago
OpenCode Go revises limits after DeepSeek price increase

OpenCode says it revised Go limits after DeepSeek raised prices. Its operator is testing hosting configurations intended to bring DeepSeek service closer to its prior price point.

NEWS1mo ago
Composio and Ante benchmark coding agent harnesses with 47%–67% success range

Composio and Ante tests reported that the same models behaved very differently by harness. DeepSeek V4 Flash ranged from 47% to 67% task success and $0.019 to $0.104 per task across harnesses.

RELEASE1mo ago
OpenRouter updates Auto Router with 30 task types and 7-day spend-based routing

OpenRouter upgraded Auto Router to classify prompts into about 30 task types, then route by anonymized 7-day spend share and cost tier. OpenRouter says the max tier beat the old router across five benchmark domains.

NEWS1mo ago
Microsoft Copilot traces report 87% of LLM calls came from agents

A Microsoft Copilot trace analysis said 87% of LLM calls came from the agent, not direct user turns. Related posts warned token use and web requests can scale far faster than human prompt counts.

NEWS1mo ago
OpenCode users reportedly average $1.14/day on DeepSeek V4 Flash

OpenCode’s thdxr said Go users spent $1.14 per day on DeepSeek V4 Flash last week. Wafer added a fast OpenRouter route, while Nous extended a 90% discount for the 0731 model.

NEWS1mo ago
DeepSeek V4 Flash benchmarks claim lower DeepSWE cost than GPT-5.6 Luna

Together says two V4 Flash attempts solved more DeepSWE tasks than one GPT-5.6 Luna attempt for roughly one-third the cost. Practitioners report Flash-0731 results vary sharply by harness and pass count.

NEWS1mo ago
Vercel details spend caps and anomaly alerts for runaway agent bills

Guillermo Rauch listed shipped safeguards including spend caps, anomaly alerts, recursion protection, billing APIs and DDoS mitigation. The post followed reports of agents looping until queues timed out.

NEWS1mo ago
DeepSeek V4 Flash benchmarks at 61.4% on ARC-AGI-2 for $0.04 per task

ARC Prize verified DeepSeek V4 Flash at 61.4% on ARC-AGI-2 for $0.04 per task. Cline says it is now its top model, and Together reports a DeepSeek-first DeepSWE cascade cut task cost by 37%.

NEWS1mo ago
Databricks reports coding-agent token spend is rising exponentially

Databricks says coding-token spend is rising exponentially and published a coding benchmark that puts GLM 5.2, Claude Opus 4.8, and GPT-5.6 Sol on the quality-per-dollar frontier. Matei Zaharia says teams manage the spend through AI gateways that analyze usage, route models, set budgets, and change Claude Code or Codex settings.

RELEASE1mo ago
DeepSeek V4 Flash adds Baseten and Together AI serving with 1M-token context

Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.

NEWS1mo ago
Cline raises free DeepSeek Flash quota 3x for coding agents

Cline tripled its free quota while Nous discounted Flash 0731 by 90% for a week and OpenHands offered free cloud use. OpenCode reported 8T Flash tokens on Aug. 1.

WORKFLOW1mo ago
Codex users route tasks across GPT-5.6 Sol, Terra, and Luna to cut token cost

Practitioners reported better Codex multi-agent runs by raising concurrency and splitting work across Sol, Terra, and Luna. One workflow sends deploy tasks to Luna Max to preserve Sol tokens.

NEWS1mo ago
DeepSeek V4 Flash benchmarks show cheaper tokens but 3x SWE-Bench task cost

New tests showed DeepSeek V4 Flash as cheaper per token and faster on some serving paths. Ramp said it cost 3x more than GPT-5.6 Luna per SWE-Bench task because it used more turns.

WORKFLOW1mo ago
Sol-advisor routes Codex tasks across GPT-5.6 Sol, Luna, and Terra

Sol-advisor routes Codex tasks through GPT-5.6 Sol, Luna, and Terra, while users compare Luna Max as a lower-cost reasoning setting. Early reports say small routing tests need larger benchmarks.

NEWS1mo ago
OpenAI cuts GPT-5.6 Luna pricing by 80%

OpenAI said GPT-5.6 Luna pricing fell 80%, while Terra fell 20%. Codex users recommended max reasoning for linting, tests, and dependency work, but cautioned against forcing Luna into subagent roles.

NEWS1mo ago
OpenAI cuts GPT-5.6 Luna API prices by 80%

OpenAI said GPT-5.6 Luna is 80% cheaper and Terra is 20% cheaper, with lower usage burn in Codex and ChatGPT Work. Sol Fast adds up to 2.5x speed at 2x price, and gateways reflected the new pricing.

NEWS2mo ago
OpenAI says GPT-5.6 Sol cuts model-serving costs by 20%

OpenAI says it used GPT-5.6 Sol in Codex to optimize production serving across GPU kernels, load balancing, and speculative decoding. The company reports a 20% end-to-end cost reduction.

NEWS2mo ago
Composio benchmarks Kimi K3 harnesses with $0.22–$2 per-task cost swing

Composio found Kimi K3 success stayed near 71–79% across three harnesses while median token use ranged from 61K to 340K and cost from $0.22 to $2 per task. Cline also shipped a Kimi K3 CLI update.

RELEASE2mo ago
Gemini API adds token budget caps for Managed Agents

Google added token budget caps and other controls for Managed Agents in the Gemini API. The release also adds sandbox hooks, cron triggers, model configuration, free-tier support, and Gemini 3.6 Flash defaults.

WORKFLOW2mo ago
Developers report under-10% unattended completion for coding agents

Robert C. Martin argued tangled code confuses agents and promoted agent-led refactoring with tests. Other reports cited high Sol spend, under-10% unattended completion, and failures once bugs or complexity appear.

NEWS2mo ago
Liang Wenfeng reportedly frames DeepSeek roadmap around scarce GPU supply

Posts quoting Liang Wenfeng said DeepSeek is targeting low positive API margins while constrained by GPU supply. They also said early-June capacity was about 20,000 H100-equivalent units and that Huawei capacity remains below Nvidia.

RELEASE2mo ago
Cursor Router adds Cost mode for Teams and Enterprise coding requests

Cursor launched Router for Teams and Enterprise with Intelligence, Balance, and Cost modes. Cursor says the router can select models per request and cut costs by 60%, with admin controls for businesses.

RELEASE2mo ago
Google ships Gemini 3.6 Flash and 3.5 Flash-Lite to serving platforms

Gemini 3.6 Flash and 3.5 Flash-Lite went live on OpenRouter, Venice, Hyperbrowser, and Google surfaces. Early benchmarks show lower token use and cost, but mixed document and coding results.

RELEASE2mo ago
Martian launches Ship beta with 50% lower-cost inference target

Martian launched Ship, a beta endpoint that sits between an app and its reference model and routes each request through cheaper paths while aiming to preserve behavior. Martian says the beta targets 50% lower cost.

NEWS2mo ago
METR introduces expenditure horizon for cost-aware agent evals

METR introduced expenditure horizon, a method that compares human and agent performance as a function of spend on continuously scored tasks. A correction estimates human labor returns near $2.5K per 1% optimization, making cost curves central to the evaluation.

NEWS2mo ago
Artificial Analysis reports Kimi K3 averages 56.4 minutes on AA-Briefcase

Artificial Analysis reports Kimi K3 averages 56.4 minutes, 83 turns, and 120k output tokens per AA-Briefcase task. Kilo also found UI-build outputs close to Claude Fable 5 at 29% of the cost.

WORKFLOW2mo ago
Engineers replace broad agent loops with scoped workflows and SWE-bench harnesses

Practitioner threads favored narrow tasks, deterministic steps, task-specific harnesses and explicit acceptance criteria over open-ended agent loops. One Codex workflow reported 58.6%-87.0% token cuts on SWE-bench task sets.

NEWS2mo ago
Kimi K3 benchmarks last at 53/67 in AlphaSignal repair harness

AlphaSignal's repair harness put Kimi K3 last at 53/67, while other tests ranked it high on DeepSWE, Vibe Code Bench, legal work, cyber, and CUDA kernels. Cost often beat Fable, but speed lagged.

NEWS2mo ago
Kimi K3 ranks #5 on Artificial Analysis as engineers dispute coding cost

Kimi K3 posted strong coding results, including rank #5 on Artificial Analysis and #3 on DeepSWE. Engineers disputed whether its lower token price offsets higher token use and slower throughput.

NEWS2mo ago
Cursor users report agents switching to costly Claude Opus, Sonnet, Fable, or API calls

Cursor users reported unintended Claude Opus, Sonnet, Fable, or API calls after selecting other settings. Reports included 22.6M-token burns, hard usage stops, and unexpected bills.

RELEASE2mo ago
Devin Fusion adds Fable 5 to cut coding-agent task costs

Cognition said Devin Fusion now uses Fable 5 and saw lower cost per task than Opus 4.8. Practitioners cited Fable-led delegation patterns that cut token use, with caveats on serial debugging.

NEWS2mo ago
Coding Agent Index ranks cheaper configs near the top

Fresh runs and charts put GPT-5.6 Sol high on SWE-Bench Pro and Design Arena, while Coding Agent Index and Amp reports emphasized cheaper strong configs. Results vary by harness, effort tier, and agent setup.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.