Skip to content
AI Primer
TOPIC27 stories

Benchmark

Stories, products, and related signals connected to this tag in Explore.

RELEASE1w ago
OpenArt Arena ranks image and video models by production task with blind pairwise judgments

OpenArt Arena uses blind pairwise judgments from creative professionals to rank image and video models for ads, film, animation, editing, and lip sync. It also publishes overall leaderboards.

WORKFLOW1w ago
A fly-brain simulation designs a buildable guitar

A 166,700-neuron fly-brain simulation placed a guitar body and headstock while scale-length math constrained the neck. Related tests used the system to draw between 23 anchor points and turn neural spikes into a 643-note performance.

RELEASE2w ago
Claude Code adds `plugin eval` command to compare plugin test runs

Claude Code’s `plugin eval` command runs test cases with and without a plugin, scores both runs, and produces terminal and HTML comparisons. Anthropic says developers can recheck skills after model releases, though evals consume tokens.

NEWS2w ago
GPT Image 2.5 reportedly improves image consistency over GPT Image 2.0

Creator and product comparisons report better realism, style transfer, character preservation, and source-image detail retention in GPT Image 2.5 than GPT Image 2.0. Repeat-regeneration tests found stronger detail retention, though darker-image results remained a caveat.

WORKFLOW2w ago
RocketLeagueBench critics question GPT-6 Astra scores versus visible game quality

A Meta Muse Spark 1.3 Max and GPT-6 Astra comparison prompted criticism that benchmark scores do not reflect visible game quality. The benchmark's creator says Rocket League's known mechanics make it useful for testing 3D replication.

RELEASE3w ago
Google releases Gemini 3.8 Flash at Gemini 3.7 Flash pricing

Google released Gemini 3.8 Flash across its developer and consumer products. Google says it matches Gemini 3.7 Flash pricing and scored 73.7% on DeepSWE 1.1.

RELEASE4w ago
MiniMax H3 Max renders 15-second 720p clips in about 16 seconds

Creator tests on Fal put MiniMax H3 Max 15-second 720p renders at roughly 16–17 seconds. A separate 480p test took 5.17 seconds and cost $0.7125.

RELEASE4w ago
JFrog releases Boost CLI, claims 13.5% coding-agent cost reduction

JFrog released the free Boost CLI, which compresses shell output before it reaches coding agents including Claude Code and Codex. Its Terminal-Bench 2.0 report claims a 13.5% cost reduction across 89 tasks.

RELEASE1mo ago
DeepSeek reportedly adds V4 Flash Vision to its API

DeepSeek's experimental V4 Flash Vision model is reportedly live on its official API. In a screenshot-only form-filling test, it completed and checked fields in 5 minutes 48 seconds but misplaced check marks and drifted in longer boxes.

RELEASE1mo ago
Alibaba releases Qwen3.8-Max with 2.4T MoE and 1M context

Alibaba released Qwen3.8-Max, described in a launch thread as a 2.4T sparse MoE with 95B active parameters, 1M context, native vision/text, and agent benchmarks. API pricing is listed at $2/$6 per 1M tokens, with open weights planned for Hugging Face.

WORKFLOW1mo ago
MiniMax H3 benchmarks local runs on RTX 3060 to 5090 GPUs

Reddit tests show MiniMax H3 running locally on GPUs from RTX 3060 to 5090. A 3060 reportedly takes about an hour for 10 seconds, while a 5090 makes 15 seconds in 84 seconds, with users also reporting motion and reference-video failures.

RELEASE1mo ago
Graft adds repo context cache for Claude Code

Graft scans a repo into linked markdown and plugs into Claude Code hooks so the agent can reuse codebase context. Its 162-run benchmark claims 46% fewer tool calls, up to 4x fewer tokens, and 60% less time.

NEWS1mo ago
MiniMax H3 adds native 2K, 15-second generation to more creator tools

TopviewAI posts cite native 2K, 15-second generations, multimodal control, and claimed pricing at 30% of Seedance 2.0. Creators also tested H3 on ComfyUI Cloud, with ComfyUI support for open weights reported.

RELEASE1mo ago
Flux 3 previews as an open-weights video model with 720p samples

Flux 3 appeared in early preview on Nous Hermes Agent as an open-weights video model, with 720p sample tests circulating. Curious Refuge ranked its early results behind Seedance 2.0 and LTX 2.3, making output quality the main launch caveat.

RELEASE2mo ago
Anthropic ships Claude Opus 5 in Claude Code and the API

Anthropic launched Claude Opus 5 for Claude Code and the API with high-effort defaults, fast mode, migration tooling, mid-conversation tool changes, and routing fallback. The model scored 43.3% on Frontier-Bench v0.1.

RELEASE2mo ago
Google releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber for cheaper agents

Google says Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber target faster, cheaper agent workloads. Early tests cite lower token use and stronger coding results.

NEWS2mo ago
Users report Kimi K3 tasks can burn 20% of a weekly $20 plan

Posts compared Kimi K3 with Fable on Rocket League-style app builds, finding stronger UI than physics and game feel. Other users said one large prompt could burn a 5-hour window and 20% of a weekly $20 plan.

NEWS2mo ago
Terra 5.6 high cuts Clawsweeper GitHub review time by ~40%, Steipete says

Steipete said he moved the Clawsweeper GitHub review bot to 5.6 Terra high and saw about 40% faster reviews at lower cost with little quality change. A follow-up framed the result as a warning that general model benchmarks may not predict issue-review performance.

NEWS2mo ago
Kimi K3 benchmarks question cost advantage as token use rises

Illscience found Kimi K3 strong on prose reasoning but costly per task because it used more tokens than expected. LLMJunky's Rocket League-style test said the model made good UI but lacked Fable-level polish.

RELEASE2mo ago
Kimi K3 gets creator tests against GPT-5.6 Sol and Fable 5

Creators posted open-weights benchmarks and tests comparing Kimi K3 with GPT-5.6 Sol and Fable 5. Demos covered UI animation, games, Seedance renders, kernel code, and reported OpenRouter rate limits.

RELEASE2mo ago
BytePlus previews Seedance 2.5 with Michael Owen football footage

BytePlus posted a Michael Owen football recreation made with Seedance 2.5 as creators treated it as a realism benchmark. Posts cite a Friday launch target and 30-second generations, pending no delay.

NEWS2mo ago
Goodside benchmarks GPT-5.6 Sol Pro on 150- and 1,025-Pokémon crossword tests

Goodside compared Claude Fable 5 Max puzzle generation with GPT-5.6 Sol Pro solving attempts. Sol solved a 150-Pokémon empty crossword but failed the 1,025-Pokémon version, with one reported success traced to the answer key.

NEWS2mo ago
Curious Refuge ranks GPT Image 2 above Meta Muse in character tests

Curious Refuge compared Meta’s Muse image model with GPT Image 2 across character sheets, infographics, posters, cinematic stills, edits, and book covers. It favored GPT Image 2 for realism and consistency.

NEWS2mo ago
Goodside tests GPT-5.6 Sol on fake handwriting and Ghost Font prompts

Goodside tested GPT-5.6 Sol and Claude Fable 5 with fake handwriting, constrained-vocabulary prompts, and Ghost Font. Sol often answered unreadable inputs, while Fable more often refused or pushed back.

NEWS2mo ago
Goodside benchmarks GPT-5.6 Sol hallucinations on binary noise

Goodside fed GPT-5.6 Sol and Claude Fable 5 binary noise and meaningless handwriting. Sol often invented hidden text, while Fable also failed on noise but more often pushed back on scribbles.

RELEASE2mo ago
GPT-5.6 Sol gets creator tests in Figma Make and Fable 5 comparisons

GPT-5.6 Sol appeared in Figma Make and creator benchmarks against Fable 5 across design, games, and video. Testers praised browser persistence and design output, while noting higher token use and mode confusion.

WORKFLOW2mo ago
Kaigani builds video-QC benchmark after Gemini and Codex frame tests

Kaigani posted generated-video QC tests where Gemini caught subtle errors and said they are deliberately glitching clips for a benchmark. The work follows earlier Codex frame-analysis tests and focuses on detecting defects in generated video.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.