Skip to content
AI Primer
TOPIC50 stories

Benchmarks

Model-level capability and performance results, including benchmark releases and score changes.

NEWS10th August
Composio and Ante benchmark coding agent harnesses with 47%–67% success range

Composio and Ante tests reported that the same models behaved very differently by harness. DeepSeek V4 Flash ranged from 47% to 67% task success and $0.019 to $0.104 per task across harnesses.

RELEASE10th August
OpenRouter updates Auto Router with 30 task types and 7-day spend-based routing

OpenRouter upgraded Auto Router to classify prompts into about 30 task types, then route by anonymized 7-day spend share and cost tier. OpenRouter says the max tier beat the old router across five benchmark domains.

NEWS8th August
DeepSeek V4 Flash benchmarks claim lower DeepSWE cost than GPT-5.6 Luna

Together says two V4 Flash attempts solved more DeepSWE tasks than one GPT-5.6 Luna attempt for roughly one-third the cost. Practitioners report Flash-0731 results vary sharply by harness and pass count.

NEWS7th August
DeepSeek V4 Flash benchmarks at 61.4% on ARC-AGI-2 for $0.04 per task

ARC Prize verified DeepSeek V4 Flash at 61.4% on ARC-AGI-2 for $0.04 per task. Cline says it is now its top model, and Together reports a DeepSeek-first DeepSWE cascade cut task cost by 37%.

NEWS7th August
Databricks reports coding-agent token spend is rising exponentially

Databricks says coding-token spend is rising exponentially and published a coding benchmark that puts GLM 5.2, Claude Opus 4.8, and GPT-5.6 Sol on the quality-per-dollar frontier. Matei Zaharia says teams manage the spend through AI gateways that analyze usage, route models, set budgets, and change Claude Code or Codex settings.

RELEASE1w ago
Qwen3.8-Max launches on OpenRouter with 1M-token context

Alibaba's Qwen3.8-Max is live on Venice and OpenRouter while open weights are still described as coming soon. Reports cite a 2.4T-parameter model with strong Vals and vision benchmark results.

RELEASE1w ago
DeepSeek V4 Flash adds Baseten and Together AI serving with 1M-token context

Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.

NEWS1w ago
DeepSeek V4 Flash benchmarks show cheaper tokens but 3x SWE-Bench task cost

New tests showed DeepSeek V4 Flash as cheaper per token and faster on some serving paths. Ramp said it cost 3x more than GPT-5.6 Luna per SWE-Bench task because it used more turns.

RELEASE1w ago
DeepSeek releases V4 Flash 0731 as MIT-licensed open weights

DeepSeek released V4 Flash 0731 with weights, a technical report, API access, 1M context, MoE routing, and low token prices. Its cited benchmarks show gains on Artificial Analysis, Terminal-Bench, Frontend Code Arena, and agent tests.

RELEASE1w ago
Epoch adds 50 unsolved problems to FrontierMath Open Problems

Epoch added significant research problems to FrontierMath Open Problems, bringing the set to 50 unsolved math problems. Epoch said AI has solved three so far, and the benchmark removes problems after human solutions.

RELEASE1w ago
Thinking Machines releases Inkling-Small 276B open-weight MoE

Thinking Machines released Inkling-Small, a 276B-parameter MoE with 12B active parameters, multimodal inputs, and a 1M-token context window. Providers added day-zero vLLM, SGLang, Modal, and gateway support.

RELEASE2w ago
Enterprise Worlds opens ITSMBench with 93 tools for enterprise agent evals

Enterprise Worlds opened ITSMBench for executable enterprise-agent tasks with persistent state, simulated users, 93 tools, and deterministic grading. The benchmark starts with IT service-management workflows.

NEWS2w ago
Agent Arena ranks Claude Opus 5 Max No. 2 across 7,000+ agent sessions

Agent Arena says Claude Opus 5 Max placed second and Opus 5 High placed third across more than 7,000 real agentic sessions. Practitioner reports remain mixed across harnesses and prompting styles.

NEWS2w ago
Composio benchmarks Kimi K3 harnesses with $0.22–$2 per-task cost swing

Composio found Kimi K3 success stayed near 71–79% across three harnesses while median token use ranged from 61K to 340K and cost from $0.22 to $2 per task. Cline also shipped a Kimi K3 CLI update.

RELEASE2w ago
OpenAI releases GPT-Transcribe API at 3.31% AA-WER and $4.50 per 1,000 minutes

OpenAI released GPT-Live-Transcribe and GPT-Transcribe in the API for streaming and offline speech recognition. The launch adds context prompts, keywords, language hints, and WER gains, with Artificial Analysis reporting GPT-Transcribe at 3.31% AA-WER and $4.50 per 1,000 audio minutes.

RELEASE2w ago
Microsoft launches MAI-Cyber-1-Flash with 95.95% CyberGym score

Microsoft said MAI-Cyber-1-Flash inside the MDASH multi-agent security harness scored 95.95% on CyberGym. The system routes harder tasks to GPT-5.4 and coordinates more than 100 specialist agents.

NEWS2w ago
Epoch says AI solved second FrontierMath open problem with Fable 5

Epoch said an AI-generated solution found a presentation for the absolute Galois group of the 2-adic numbers, marking the second FrontierMath open problem it says AI solved. The result was elicited with Fable 5 and also GPT-5.5 Pro, making it a benchmark milestone rather than a product release.

NEWS2w ago
Developers report Claude Opus 5 reliability tradeoffs despite ProgramBench win

Claude Opus 5 combined benchmark wins with coding-agent reliability complaints. It led ProgramBench and Arena frontend rankings, while users reported overbuilding, missed intent, and switches back to Fable.

NEWS2w ago
SOOFI revises report after GPQA removal, drawing new eval-leakage criticism

Julius Jitsev said SOOFI removed GPQA and its capability index after feedback but still compared against Nemotron 3 Nano using benchmarks seen in training. He argued the remaining English and German scores are compromised.

NEWS2w ago
Claude Opus 5 ranks first on LLM Debate, DeepSWE, and other public benchmarks

New benchmark posts put Claude Opus 5 first on LLM Debate, a short-story test, Extended NYT Connections, and DeepSWE. Developers also reported over-editing and long-session failures, making private evals a recurring caveat.

RELEASE2w ago
Anthropic ships Claude Opus 5 to paid plans and API at Opus 4.8 price

Anthropic released Claude Opus 5 with Fast Mode on paid plans and the API at the Opus 4.8 price. Benchmarks from ARC Prize, Artificial Analysis, Vals AI, and tool vendors put it near or ahead of Fable 5 on several agent and coding tests.

RELEASE3w ago
Firecrawl releases /search endpoint with agent-ready excerpts

Firecrawl introduced a new /search system that returns relevant excerpts instead of full pages for agent workflows. The launch claims 94.7% on SimpleQA and roughly 10x fewer tokens than full-page processing.

RELEASE3w ago
Google ships Gemini 3.6 Flash and 3.5 Flash-Lite to serving platforms

Gemini 3.6 Flash and 3.5 Flash-Lite went live on OpenRouter, Venice, Hyperbrowser, and Google surfaces. Early benchmarks show lower token use and cost, but mixed document and coding results.

RELEASE3w ago
Poolside releases Laguna S 2.1 as 118B open-weight coding model

Poolside released Laguna S 2.1, a 118B-parameter open-weight MoE with 1M context and SGLang day-zero support. Poolside and partners cite SWE-bench, Terminal-Bench, and local-agent tests.

NEWS3w ago
METR introduces expenditure horizon for cost-aware agent evals

METR introduced expenditure horizon, a method that compares human and agent performance as a function of spend on continuously scored tasks. A correction estimates human labor returns near $2.5K per 1% optimization, making cost curves central to the evaluation.

NEWS3w ago
Artificial Analysis reports Kimi K3 averages 56.4 minutes on AA-Briefcase

Artificial Analysis reports Kimi K3 averages 56.4 minutes, 83 turns, and 120k output tokens per AA-Briefcase task. Kilo also found UI-build outputs close to Claude Fable 5 at 29% of the cost.

RELEASE3w ago
Alibaba opens Qwen 3.8 Max Preview testing across Cloud, Qwen Chat, Qoder and web

Alibaba opened Qwen 3.8 Max Preview on Alibaba Cloud, Qwen Chat, Qoder and the web, describing it as a 2.4T model headed for open weights. Early testers praised vision results but disputed coding claims.

NEWS3w ago
Kimi K3 ranks No. 1 on Arena Frontend Code leaderboard

Posts put Kimi K3 first on Arena's Frontend Code leaderboard but 4.37-5.29 months behind U.S. frontier models in one public estimate. Other evidence cited strong DeepSWE cost-performance and cybersecurity results.

NEWS3w ago
Kimi K3 benchmarks last at 53/67 in AlphaSignal repair harness

AlphaSignal's repair harness put Kimi K3 last at 53/67, while other tests ranked it high on DeepSWE, Vibe Code Bench, legal work, cyber, and CUDA kernels. Cost often beat Fable, but speed lagged.

NEWS3w ago
Code Arena users report Kaleb shows Qwen-like token quirks

Multiple testers reported Kaleb and torenia-alpha appearing in Code Arena/LMArena, and several pointed to Qwen-like token quirks in Kaleb. Follow-up tests described strong 3D output and a late-2025 or early-2026 cutoff, but the model identity remains unconfirmed.

NEWS3w ago
Kimi K3 ranks #5 on Artificial Analysis as engineers dispute coding cost

Kimi K3 posted strong coding results, including rank #5 on Artificial Analysis and #3 on DeepSWE. Engineers disputed whether its lower token price offsets higher token use and slower throughput.

NEWS3w ago
Posts claim GPT-5.6 Sol beats Mythos 5 on UK AISI and CyberGym tasks

Posts citing UK AISI and CyberGym said GPT-5.6 Sol beat Mythos 5 on narrow cyber tasks and The Last Ones. Greg Brockman separately invited defenders to test it on real systems.

RELEASE3w ago
Moonshot launches Kimi K3 with 2.8T parameters and 1M context

Moonshot launched Kimi K3 in Kimi products and API with 1M context, native multimodality, KDA/AttnRes, and weights promised by July 27. Benchmarks place it near frontier systems, but testers cite slow serving and usability caveats.

NEWS4w ago
Users report Kivine on LMArena may be a Kimi K3 preview

Testers say Kivine identifies with Moonshot/Kimi and produces strong frontend, coding, and spatial demos. Moonshot also teased Kimi K3, but the Arena claims remain unofficial.

RELEASE4w ago
Perplexity releases WANDR benchmark with 500 deep-research tasks

Perplexity released WANDR, its internal benchmark for deep and wide research in Computer. The dataset has 500 tasks, 170,495 source-backed records and production-derived use cases.

NEWS4w ago
Meta says its model scored 30/30 on Asian Physics Olympiad theory exam

Meta said a model scored 30/30 on the APhO theoretical exam. Team posts described data curation, training and live participation, while public posts questioned which model was evaluated.

NEWS4w ago
Muse Spark 1.1 benchmarks at 863 Elo on AA-Briefcase

Posts put Muse Spark 1.1 ahead of GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam, third on Debate Benchmark, and at 863 Elo on AA-Briefcase. The same posts noted weaker presentation quality.

RELEASE4w ago
Skyfall AI launches Morpheus benchmark for persistent-world agents

Morpheus gives models persistent simulation environments where rules, objectives, and consequences shift without resets. Early reports said frontier models leaned on pretraining heuristics.

RELEASE4w ago
German consortium releases SOOFI base model trained on 27T tokens

A German consortium released the small SOOFI sovereign base model trained on 27T tokens. Analysts said it reuses Nemotron 3 Nano architecture and many hyperparameters with a changed data mix, and benchmarks drew criticism for overstating capability versus Qwen and Nemotron.

NEWS4w ago
Coding Agent Index ranks cheaper configs near the top

Fresh runs and charts put GPT-5.6 Sol high on SWE-Bench Pro and Design Arena, while Coding Agent Index and Amp reports emphasized cheaper strong configs. Results vary by harness, effort tier, and agent setup.

NEWS4w ago
Devin exposes Fusion router with claimed 35% lower frontier-model cost

Devin exposed routers such as Fusion, claiming frontier performance at 35% lower cost, while Databricks argued for smart routing in agent harnesses. New charts put Grok 4.5 and Muse Spark 1.1 near the coding-cost frontier.

RELEASE4w ago
LingBot 2.0 releases open weights for robot-action and world models

LingBot 2.0 released code and weights for a real-time world model and robot-action models. The VLA maps 20 robot body configurations into a 55D action format and filters 90,000 raw robot hours to 50,000 training hours.

NEWS4w ago
GPT-5.6 Sol ranks near top of DeepSWE and coding evals at lower reported cost

New benchmark posts put GPT-5.6 Sol at or near the top of DeepSWE and several coding/context evals. Cost reports placed Luna on the efficiency frontier, while Amp said replacing Opus with GPT-5.6 cut its average model costs ~50%.

NEWS4w ago
Muse Spark 1.1 benchmarks near frontier agent models at $1.25/M input

Meta and third-party benchmark posts put Muse Spark 1.1 near frontier coding and agent models at $1.25/M input and $4.25/M output. Results included Vals AI agent tasks, Code Arena Frontend #9, and an AA Coding Agent Index score of 69.

RELEASE4w ago
Perplexity adds Grok 4.5 as Computer orchestrator with 0.328 WANDR score

Perplexity added Grok 4.5 as an orchestrator model in Computer for Pro, Max, and Enterprise users. Perplexity reported a WANDR score of 0.328 at $4.76 per trial, while outside security-review and canvas-task tests put it close to GPT-5.6 Sol on cost or token use.

NEWS4w ago
Early benchmarks rank GPT-5.6 Sol near Fable 5 at lower cost

ARC Prize, Artificial Analysis, CursorBench, and other tests reported strong GPT-5.6 Sol results, especially in coding-agent tasks. Results were uneven, with smaller gains in document parsing and some UI or puzzle evals.

RELEASE4w ago
Genspark and OpenClaw add Grok 4.5 for coding-agent workflows

Genspark and OpenClaw added Grok 4.5 after xAI's launch, extending the model into more coding-agent workflows. Follow-up evidence covered AA-Briefcase and Terminal-Bench results, a Composio credential-audit run, and SuperGrok usage-meter reports.

NEWS1mo ago
Databricks benchmarks coding agents on internal codebase tasks

Databricks published an internal coding-agent benchmark using tasks from its codebase. OpenAI, Anthropic, and GLM-5.2 models landed on its Pareto frontier, and the company argues teams should optimize cost per task rather than per token.

RELEASE1mo ago
Grok 4.5 launches for coding agents at $2/M input and $6/M output

SpaceXAI launched Grok 4.5 in Cursor and several agent tools with $2/M input and $6/M output pricing. Early evals place it near frontier coding models, with 51% on AutomationBench-AA.

NEWS1mo ago
OpenAI audits SWE-Bench Pro and finds 30% of public tasks broken

OpenAI audited SWE-Bench Pro and found 30% of public tasks were broken. It retracted its earlier recommendation to use the benchmark as a leading coding eval.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.