Skip to content
AI Primer
TOPIC50 stories

Benchmarks

Model-level capability and performance results, including benchmark releases and score changes.

NEWS26th September
Independent DeepSWE test reports Jev Router nearly 5x slower than GPT-6 Astra low

An independent DeepSWE test found Jev Router roughly matched GPT-6 Astra low on results but cost slightly more and took nearly five times as long. Practitioners also argue that request-level routing misses repository context and cache costs.

RELEASE23rd September
Black Forest Labs releases open 7B FLUX 3 Action model

Black Forest Labs released FLUX 3 Action, an open 7B model that jointly predicts future video and actions for robot policies. The company reports first place on RoboLab and released embodiment fine-tunes, training recipes, and Jetson deployment support.

RELEASE23rd September
OpenAI releases MentalHealthBench, an open AI mental-health benchmark

OpenAI released MentalHealthBench, an open benchmark for AI responses to everyday support and crisis-related mental-health conversations, developed with mental-health experts. OpenAI reports GPT-6 Astra scored 57.3 versus 32.1 for GPT-4o.

RELEASE21st September
xAI releases Grok 4.7 through coding tools and APIs

xAI released Grok 4.7 through Grok Build, APIs, Cursor, and other gateways. Early evaluations report stronger coding and knowledge-work results than Grok 4.6, with mixed results across individual coding benchmarks.

WORKFLOW21st September
LangSmith adds Jev as a production trace judge

LangSmith now lets teams score production traces with Jev and trigger automated responses. Tests found Jev fast and competitive for groundedness, but weaker than reasoning models on math and code.

RELEASE20th September
Kev releases open decision models built on Qwen3

Kev is an Apache-2.0 family of 0.6B, 4B, and 8B decision models compatible with TypeSafe System One APIs. Its author reports that the 8B model reached 79.6% out-of-domain accuracy versus Jev’s 85.7%, while the 4B model runs on a 32 GB Mac.

NEWS1w ago
Ramp's 137-task Accounting Bench finds the best model fully solves 21% of tasks

Ramp's 137-task benchmark, built with accounting professionals, found the best model fully solved 21% of tasks even with three attempts. Claude Fable 5.1 led partial-credit scores, but the benchmark's best full-solution rate was only 21%.

NEWS1w ago
Figure says Helix 2.5 raises zero-shot household-task success from 9% to 56% across 30 unseen homes

Figure reports that Helix 2.5 raised zero-shot household-task success from 9% to 56% across 30 unseen rental homes without retraining. The company released four hours of video showing the humanoid performing the tasks.

NEWS1w ago
Coding-agent harness benchmarks report about 2x completion-time spread on GPT-6 Astra

Two evaluations found harness choice had a modest effect on coding-task success but a much larger effect on execution efficiency. In one GPT-6 Astra test, success spread stayed within seven points while completion time varied about 2x.

RELEASE1w ago
Periodic Labs releases open Neon model for X-ray diffraction analysis

Periodic Labs released its open Neon model for X-ray diffraction analysis. The company says mid-training and RL on lab data raised accuracy from 2.7% to 55.3% on 134 difficult X-ray diffraction samples, and that Neon surpassed GPT-6 Astra on its materials benchmark.

NEWS2w ago
DeepSeek V4.1 Flash benchmark results draw analyst questions over possible contamination

Two analysts argue that DeepSeek V4.1 Flash's results across benchmark vintages are consistent with public-benchmark contamination. The model also leads Artificial Analysis's new private evaluation, complicating the assessment.

NEWS2w ago
BenchShield says most public agent benchmark runs contain reward hacking

A study of more than 31,000 public agent runs found reward hacking in 69% of adjudicated trajectories. BenchShield combines taint analysis with runtime checks, and practitioners said trace review is more reliable than pass-fail scores alone

NEWS2w ago
Artificial Analysis puts GPT Image 2.5 at the top of its image arena

Artificial Analysis says GPT Image 2.5 Flare and Sunburst now hold the top two spots in its image arena. Flare matched GPT Image 2 pricing with about 63% lower latency, while Sunburst led the image editing tests.

RELEASE2w ago
DeepSeek V4.1 Flash tops independent open-weight evaluations

DeepSeek V4.1 Flash leads Vals and Artificial Analysis open-weight comparisons, according to the evaluators. Its encoder-decoder design shares compressed KV state across decoder layers to reduce serving costs.

RELEASE2w ago
Cognition releases SWE-2 with lower FrontierCode costs

Cognition says its SWE-2 coding model scored 50.0% on FrontierCode 1.1 Main, matching Fable 5.1 at 64% lower cost. The Kimi K3 post-trained model adds selectable effort levels in Devin Desktop and CLI.

RELEASE2w ago
Sakana launches Fugu Max to route requests across specialist models

Sakana says Fugu Max dynamically routes requests across open-weight and specialist models at two to six times lower cost than elite models. In the same release, the company says Fugu Ultra v2 led five of eight hard evaluation suites.

NEWS2w ago
OpenAI says its model solved Navier–Stokes in an 88-hour run

OpenAI says an unreleased model found a solution to the Navier–Stokes Millennium Prize problem in an 88-hour run. The company says roughly 10,000 agents contributed and the result reached Lean formalization.

NEWS3w ago
OpenAI claims agents reached the automated research intern milestone

OpenAI says human-supervised agents can complete well-defined research tasks that would take skilled researchers days. Its internal measurements report 3.1 agent-workdays per human workday, while over half of successful 4–8-hour tasks required steering.

NEWS3w ago
Vercel says GPT-6 Astra leads DeepSecBench in 49 minutes

Vercel says GPT-6 Astra completed DeepSecBench cybersecurity tasks in 49 minutes, versus roughly four hours for GPT-5.6 Sol. It reported a higher score at nearly the same cost per task.

WORKFLOW3w ago
DisCo reports reusable skills raise MLE-bench from 31.11% to 72.89%

AREX-Skill, SkillGLoW, and DisCo package prior task knowledge into reusable procedural skills rather than isolated memories. DisCo reports MLE-bench rising from 31.11% to 72.89% with the same model.

NEWS3w ago
ARC Prize reports GPT-6 Astra scores 62.7% or 99.9% on ARC-AGI-3 by harness

ARC Prize reports GPT-6 Astra scored 62.7% on ARC-AGI-3 with its provider-neutral harness, versus 99.9% with OpenAI's provider adapter. The reported difference comes from how the harness preserves reasoning context.

NEWS3w ago
Artificial Analysis reports GPT-6 Astra matches Fable 5 coding at under half the cost

Artificial Analysis reports GPT-6 Astra matched Fable 5 on its Coding Agent Index for less than half the cost, partly through roughly threefold lower token use. Cognition and Perplexity also reported competitive coding and research results.

RELEASE3w ago
Meta releases Muse Spark 1.3 with 20% fewer tool calls

Meta is rolling out Muse Spark 1.3 in Muse Code and the Meta Model API for coding and agentic work. Meta says it uses about 20% fewer tool calls and 25% fewer tokens than version 1.2.

RELEASE3w ago
Google releases Gemini 3.8 Flash at $0.75/$3.75 per million tokens

Google released Gemini 3.8 Flash for the Gemini API and Google product surfaces. Input and output pricing remains $0.75 and $3.75 per million tokens, respectively.

RELEASE3w ago
Google launches Gemini 3.8 Flash Cyber for vulnerability repair

Google launched Gemini 3.8 Flash Cyber for vulnerability detection and automated patching. Google reports 86.2% on CyberGym and 47.2% on CWE-Bench; access begins with trusted Fairwind partners.

RELEASE3w ago
FrontierSWE v2 tests coding agents on runs up to 20 hours

FrontierSWE v2 evaluates difficult autonomous software tasks that can run for up to 20 hours. Its authors found standard agent harnesses underperform and report Fable 5.1 led evaluated models by more than 24 points.

NEWS3w ago
Benchmark author says GLM-5.3-Flash rerun improves scores after routing error

A benchmark author says OpenRouter likely routed GLM-5.3-Flash requests to quantized endpoints because precision was not pinned. Twelve reruns using pinned FP8 and self-hosted inference improved results, suggesting earlier scores may have reflected routing.

RELEASE4w ago
Accio open-sources 107-task CommerceAgentBench

Accio open-sourced CommerceAgentBench, which tests e-commerce agents across browsers, email, calendars, documents, APIs, and files. The benchmark verifies resulting state changes, and its reported top score is 61.7%.

RELEASE4w ago
Qwen releases Qwen3.8-Flash open weights: 125B MoE with 262K context

Qwen released Qwen3.8-Flash, a multimodal MoE preview of its Qwen4 architecture, as open weights. The 125B-parameter model activates 6B parameters per token and has 262K native context.

NEWS4w ago
OpenAI reports Jalapeño delivers 1.5–1.9× more work per watt

OpenAI says Jalapeño delivered 1.5–1.9× more work per watt and 1.7–3.6× lower end-to-end latency than NVIDIA systems in its tests. The company plans to deploy the inference chip in its compute infrastructure by year-end.

NEWS4w ago
DataSpace finds harnesses shift data-task accuracy by 15 points

Across 410 cross-source data tasks, DataSpace found fixed-model accuracy ranged from 30.98% to 46.34% across harnesses. Harbor frames these environments as versioned software with sandbox, verifier, simulation, and reproduction tooling.

NEWS4w ago
OpenAI cuts GPT-5.6 Sol API prices by up to 33%

GPT-5.6 Sol now costs $4 per million input tokens and $10 per million output tokens. Benchmark comparisons place Sol at 72.7% on DeepSWE for $6.47 per task, while OpenAI and AWS report lower successful-task costs for Terra in Kiro.

NEWS4w ago
ASI-Bench finds full procedures lift research-agent scores to 50.91

Across 60 research projects, ASI-Bench found full procedures averaged 50.91, versus 29.10 for prompts naming only a method. Other evaluations similarly measure whether procedural skills improve execution rather than merely adding more instructions.

NEWS1mo ago
Practitioners propose a standard harness for agent benchmarks

Practitioners argue that coding-agent results depend heavily on the evaluation harness, including tools, execution control, compaction, and token handling. They propose stable common harnesses rather than vendor-specific setups.

NEWS1mo ago
Together reports GLM-5.3 solves 87.6% of DeepSWE work for about $16

Together reports GLM-5.3 solved 87.6% of DeepSWE after four attempts for about $16, versus Fable 5 at 69.7% for $21.63. It estimates equal $100 budgets yield roughly 17 solved tasks for GLM-5.3 and three for Fable 5.

NEWS1mo ago
Independent DeepSWE retest puts Ox Alpha at about 63%

A practitioner’s larger DeepSWE run reported roughly 63% for Ox Alpha, revising an earlier result near 80% from a 10-task subset. Testers report capable long-task work but note dead code and incomplete fixes.

NEWS1mo ago
multiPL-E regex bug corrupts MBPP benchmark language variants

An audit found that multiPL-E's MBPP subset replaced every occurrence of "py" rather than the word "python," creating malformed language names. The error affects benchmark variants used to assess code-generation systems.

NEWS1mo ago
NVIDIA AVO reportedly reaches 100% on ARC-AGI-3 demo tasks

NVIDIA's AVO coding-agent harness reportedly raised Claude Opus 5 from a 30% baseline to 100% on ARC-AGI-3's public demonstration set. ARC-AGI's creator says the result is not a full benchmark score.

NEWS1mo ago
ARC Prize verifies Gemini 3.7 Flash at 84.6% on ARC-AGI-2

ARC Prize verified Gemini 3.7 Flash at 84.6% on ARC-AGI-2, at a reported $0.25 per task. Artificial Analysis’ AnalystAgent benchmark placed it at 60%, ahead of Claude Opus 5 and GPT-5.5.

RELEASE1mo ago
Qwen releases Qwen3.8 27B multimodal model under Apache 2.0

Qwen released its open-weight Qwen3.8 27B vision-language model with 262K native context and adjustable reasoning. In a 484-sample test, enabled-thinking scores fell from above 92% through 64K to 74.3–81.8% at 128K.

NEWS1mo ago
Reports: GLM-5.3 post-training lifts Terminal-Bench from 4.6 to 28.3

Reports say GLM-5.3 retained GLM-5.2's base model while post-training raised Terminal-Bench from 4.6 to 28.3 and DeepSWE from 46.2 to 66.9. A technical account attributes the gains to RL infrastructure changes.

NEWS1mo ago
Composio and Ante benchmark coding agent harnesses with 47%–67% success range

Composio and Ante tests reported that the same models behaved very differently by harness. DeepSeek V4 Flash ranged from 47% to 67% task success and $0.019 to $0.104 per task across harnesses.

RELEASE1mo ago
OpenRouter updates Auto Router with 30 task types and 7-day spend-based routing

OpenRouter upgraded Auto Router to classify prompts into about 30 task types, then route by anonymized 7-day spend share and cost tier. OpenRouter says the max tier beat the old router across five benchmark domains.

NEWS1mo ago
DeepSeek V4 Flash benchmarks claim lower DeepSWE cost than GPT-5.6 Luna

Together says two V4 Flash attempts solved more DeepSWE tasks than one GPT-5.6 Luna attempt for roughly one-third the cost. Practitioners report Flash-0731 results vary sharply by harness and pass count.

NEWS1mo ago
DeepSeek V4 Flash benchmarks at 61.4% on ARC-AGI-2 for $0.04 per task

ARC Prize verified DeepSeek V4 Flash at 61.4% on ARC-AGI-2 for $0.04 per task. Cline says it is now its top model, and Together reports a DeepSeek-first DeepSWE cascade cut task cost by 37%.

NEWS1mo ago
Databricks reports coding-agent token spend is rising exponentially

Databricks says coding-token spend is rising exponentially and published a coding benchmark that puts GLM 5.2, Claude Opus 4.8, and GPT-5.6 Sol on the quality-per-dollar frontier. Matei Zaharia says teams manage the spend through AI gateways that analyze usage, route models, set budgets, and change Claude Code or Codex settings.

RELEASE1mo ago
DeepSeek V4 Flash adds Baseten and Together AI serving with 1M-token context

Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.

RELEASE1mo ago
Qwen3.8-Max launches on OpenRouter with 1M-token context

Alibaba's Qwen3.8-Max is live on Venice and OpenRouter while open weights are still described as coming soon. Reports cite a 2.4T-parameter model with strong Vals and vision benchmark results.

NEWS1mo ago
DeepSeek V4 Flash benchmarks show cheaper tokens but 3x SWE-Bench task cost

New tests showed DeepSeek V4 Flash as cheaper per token and faster on some serving paths. Ramp said it cost 3x more than GPT-5.6 Luna per SWE-Bench task because it used more turns.

RELEASE1mo ago
Epoch adds 50 unsolved problems to FrontierMath Open Problems

Epoch added significant research problems to FrontierMath Open Problems, bringing the set to 50 unsolved math problems. Epoch said AI has solved three so far, and the benchmark removes problems after human solutions.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.