Benchmarks
Model-level capability and performance results, including benchmark releases and score changes.
Stories
Filter storiesComposio and Ante tests reported that the same models behaved very differently by harness. DeepSeek V4 Flash ranged from 47% to 67% task success and $0.019 to $0.104 per task across harnesses.
OpenRouter upgraded Auto Router to classify prompts into about 30 task types, then route by anonymized 7-day spend share and cost tier. OpenRouter says the max tier beat the old router across five benchmark domains.
Together says two V4 Flash attempts solved more DeepSWE tasks than one GPT-5.6 Luna attempt for roughly one-third the cost. Practitioners report Flash-0731 results vary sharply by harness and pass count.
ARC Prize verified DeepSeek V4 Flash at 61.4% on ARC-AGI-2 for $0.04 per task. Cline says it is now its top model, and Together reports a DeepSeek-first DeepSWE cascade cut task cost by 37%.
Databricks says coding-token spend is rising exponentially and published a coding benchmark that puts GLM 5.2, Claude Opus 4.8, and GPT-5.6 Sol on the quality-per-dollar frontier. Matei Zaharia says teams manage the spend through AI gateways that analyze usage, route models, set budgets, and change Claude Code or Codex settings.
Alibaba's Qwen3.8-Max is live on Venice and OpenRouter while open weights are still described as coming soon. Reports cite a 2.4T-parameter model with strong Vals and vision benchmark results.
Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.
New tests showed DeepSeek V4 Flash as cheaper per token and faster on some serving paths. Ramp said it cost 3x more than GPT-5.6 Luna per SWE-Bench task because it used more turns.
DeepSeek released V4 Flash 0731 with weights, a technical report, API access, 1M context, MoE routing, and low token prices. Its cited benchmarks show gains on Artificial Analysis, Terminal-Bench, Frontend Code Arena, and agent tests.
Epoch added significant research problems to FrontierMath Open Problems, bringing the set to 50 unsolved math problems. Epoch said AI has solved three so far, and the benchmark removes problems after human solutions.
Thinking Machines released Inkling-Small, a 276B-parameter MoE with 12B active parameters, multimodal inputs, and a 1M-token context window. Providers added day-zero vLLM, SGLang, Modal, and gateway support.
Enterprise Worlds opened ITSMBench for executable enterprise-agent tasks with persistent state, simulated users, 93 tools, and deterministic grading. The benchmark starts with IT service-management workflows.
Agent Arena says Claude Opus 5 Max placed second and Opus 5 High placed third across more than 7,000 real agentic sessions. Practitioner reports remain mixed across harnesses and prompting styles.
Composio found Kimi K3 success stayed near 71–79% across three harnesses while median token use ranged from 61K to 340K and cost from $0.22 to $2 per task. Cline also shipped a Kimi K3 CLI update.
OpenAI released GPT-Live-Transcribe and GPT-Transcribe in the API for streaming and offline speech recognition. The launch adds context prompts, keywords, language hints, and WER gains, with Artificial Analysis reporting GPT-Transcribe at 3.31% AA-WER and $4.50 per 1,000 audio minutes.
Microsoft said MAI-Cyber-1-Flash inside the MDASH multi-agent security harness scored 95.95% on CyberGym. The system routes harder tasks to GPT-5.4 and coordinates more than 100 specialist agents.
Epoch said an AI-generated solution found a presentation for the absolute Galois group of the 2-adic numbers, marking the second FrontierMath open problem it says AI solved. The result was elicited with Fable 5 and also GPT-5.5 Pro, making it a benchmark milestone rather than a product release.
Claude Opus 5 combined benchmark wins with coding-agent reliability complaints. It led ProgramBench and Arena frontend rankings, while users reported overbuilding, missed intent, and switches back to Fable.
Julius Jitsev said SOOFI removed GPQA and its capability index after feedback but still compared against Nemotron 3 Nano using benchmarks seen in training. He argued the remaining English and German scores are compromised.
New benchmark posts put Claude Opus 5 first on LLM Debate, a short-story test, Extended NYT Connections, and DeepSWE. Developers also reported over-editing and long-session failures, making private evals a recurring caveat.
Anthropic released Claude Opus 5 with Fast Mode on paid plans and the API at the Opus 4.8 price. Benchmarks from ARC Prize, Artificial Analysis, Vals AI, and tool vendors put it near or ahead of Fable 5 on several agent and coding tests.
Firecrawl introduced a new /search system that returns relevant excerpts instead of full pages for agent workflows. The launch claims 94.7% on SimpleQA and roughly 10x fewer tokens than full-page processing.
Gemini 3.6 Flash and 3.5 Flash-Lite went live on OpenRouter, Venice, Hyperbrowser, and Google surfaces. Early benchmarks show lower token use and cost, but mixed document and coding results.
Poolside released Laguna S 2.1, a 118B-parameter open-weight MoE with 1M context and SGLang day-zero support. Poolside and partners cite SWE-bench, Terminal-Bench, and local-agent tests.
METR introduced expenditure horizon, a method that compares human and agent performance as a function of spend on continuously scored tasks. A correction estimates human labor returns near $2.5K per 1% optimization, making cost curves central to the evaluation.
Artificial Analysis reports Kimi K3 averages 56.4 minutes, 83 turns, and 120k output tokens per AA-Briefcase task. Kilo also found UI-build outputs close to Claude Fable 5 at 29% of the cost.
Alibaba opened Qwen 3.8 Max Preview on Alibaba Cloud, Qwen Chat, Qoder and the web, describing it as a 2.4T model headed for open weights. Early testers praised vision results but disputed coding claims.
Posts put Kimi K3 first on Arena's Frontend Code leaderboard but 4.37-5.29 months behind U.S. frontier models in one public estimate. Other evidence cited strong DeepSWE cost-performance and cybersecurity results.
AlphaSignal's repair harness put Kimi K3 last at 53/67, while other tests ranked it high on DeepSWE, Vibe Code Bench, legal work, cyber, and CUDA kernels. Cost often beat Fable, but speed lagged.
Multiple testers reported Kaleb and torenia-alpha appearing in Code Arena/LMArena, and several pointed to Qwen-like token quirks in Kaleb. Follow-up tests described strong 3D output and a late-2025 or early-2026 cutoff, but the model identity remains unconfirmed.
Kimi K3 posted strong coding results, including rank #5 on Artificial Analysis and #3 on DeepSWE. Engineers disputed whether its lower token price offsets higher token use and slower throughput.
Posts citing UK AISI and CyberGym said GPT-5.6 Sol beat Mythos 5 on narrow cyber tasks and The Last Ones. Greg Brockman separately invited defenders to test it on real systems.
Moonshot launched Kimi K3 in Kimi products and API with 1M context, native multimodality, KDA/AttnRes, and weights promised by July 27. Benchmarks place it near frontier systems, but testers cite slow serving and usability caveats.
Testers say Kivine identifies with Moonshot/Kimi and produces strong frontend, coding, and spatial demos. Moonshot also teased Kimi K3, but the Arena claims remain unofficial.
Perplexity released WANDR, its internal benchmark for deep and wide research in Computer. The dataset has 500 tasks, 170,495 source-backed records and production-derived use cases.
Meta said a model scored 30/30 on the APhO theoretical exam. Team posts described data curation, training and live participation, while public posts questioned which model was evaluated.
Posts put Muse Spark 1.1 ahead of GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam, third on Debate Benchmark, and at 863 Elo on AA-Briefcase. The same posts noted weaker presentation quality.
Morpheus gives models persistent simulation environments where rules, objectives, and consequences shift without resets. Early reports said frontier models leaned on pretraining heuristics.
A German consortium released the small SOOFI sovereign base model trained on 27T tokens. Analysts said it reuses Nemotron 3 Nano architecture and many hyperparameters with a changed data mix, and benchmarks drew criticism for overstating capability versus Qwen and Nemotron.
Fresh runs and charts put GPT-5.6 Sol high on SWE-Bench Pro and Design Arena, while Coding Agent Index and Amp reports emphasized cheaper strong configs. Results vary by harness, effort tier, and agent setup.
Devin exposed routers such as Fusion, claiming frontier performance at 35% lower cost, while Databricks argued for smart routing in agent harnesses. New charts put Grok 4.5 and Muse Spark 1.1 near the coding-cost frontier.
LingBot 2.0 released code and weights for a real-time world model and robot-action models. The VLA maps 20 robot body configurations into a 55D action format and filters 90,000 raw robot hours to 50,000 training hours.
New benchmark posts put GPT-5.6 Sol at or near the top of DeepSWE and several coding/context evals. Cost reports placed Luna on the efficiency frontier, while Amp said replacing Opus with GPT-5.6 cut its average model costs ~50%.
Meta and third-party benchmark posts put Muse Spark 1.1 near frontier coding and agent models at $1.25/M input and $4.25/M output. Results included Vals AI agent tasks, Code Arena Frontend #9, and an AA Coding Agent Index score of 69.
Perplexity added Grok 4.5 as an orchestrator model in Computer for Pro, Max, and Enterprise users. Perplexity reported a WANDR score of 0.328 at $4.76 per trial, while outside security-review and canvas-task tests put it close to GPT-5.6 Sol on cost or token use.
ARC Prize, Artificial Analysis, CursorBench, and other tests reported strong GPT-5.6 Sol results, especially in coding-agent tasks. Results were uneven, with smaller gains in document parsing and some UI or puzzle evals.
Genspark and OpenClaw added Grok 4.5 after xAI's launch, extending the model into more coding-agent workflows. Follow-up evidence covered AA-Briefcase and Terminal-Bench results, a Composio credential-audit run, and SuperGrok usage-meter reports.
Databricks published an internal coding-agent benchmark using tasks from its codebase. OpenAI, Anthropic, and GLM-5.2 models landed on its Pareto frontier, and the company argues teams should optimize cost per task rather than per token.
SpaceXAI launched Grok 4.5 in Cursor and several agent tools with $2/M input and $6/M output pricing. Early evals place it near frontier coding models, with 51% on AutomationBench-AA.
OpenAI audited SWE-Bench Pro and found 30% of public tasks were broken. It retracted its earlier recommendation to use the benchmark as a leading coding eval.