Skip to content
AI Primer
TOPIC18 stories

Benchmark

Stories, products, and related signals connected to this tag in Explore.

RELEASE9th August
Alibaba releases Qwen3.8-Max with 2.4T MoE and 1M context

Alibaba released Qwen3.8-Max, described in a launch thread as a 2.4T sparse MoE with 95B active parameters, 1M context, native vision/text, and agent benchmarks. API pricing is listed at $2/$6 per 1M tokens, with open weights planned for Hugging Face.

WORKFLOW9th August
MiniMax H3 benchmarks local runs on RTX 3060 to 5090 GPUs

Reddit tests show MiniMax H3 running locally on GPUs from RTX 3060 to 5090. A 3060 reportedly takes about an hour for 10 seconds, while a 5090 makes 15 seconds in 84 seconds, with users also reporting motion and reference-video failures.

RELEASE1w ago
Graft adds repo context cache for Claude Code

Graft scans a repo into linked markdown and plugs into Claude Code hooks so the agent can reuse codebase context. Its 162-run benchmark claims 46% fewer tool calls, up to 4x fewer tokens, and 60% less time.

NEWS1w ago
MiniMax H3 adds native 2K, 15-second generation to more creator tools

TopviewAI posts cite native 2K, 15-second generations, multimodal control, and claimed pricing at 30% of Seedance 2.0. Creators also tested H3 on ComfyUI Cloud, with ComfyUI support for open weights reported.

RELEASE1w ago
Flux 3 previews as an open-weights video model with 720p samples

Flux 3 appeared in early preview on Nous Hermes Agent as an open-weights video model, with 720p sample tests circulating. Curious Refuge ranked its early results behind Seedance 2.0 and LTX 2.3, making output quality the main launch caveat.

RELEASE2w ago
Anthropic ships Claude Opus 5 in Claude Code and the API

Anthropic launched Claude Opus 5 for Claude Code and the API with high-effort defaults, fast mode, migration tooling, mid-conversation tool changes, and routing fallback. The model scored 43.3% on Frontier-Bench v0.1.

RELEASE3w ago
Google releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber for cheaper agents

Google says Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber target faster, cheaper agent workloads. Early tests cite lower token use and stronger coding results.

NEWS3w ago
Users report Kimi K3 tasks can burn 20% of a weekly $20 plan

Posts compared Kimi K3 with Fable on Rocket League-style app builds, finding stronger UI than physics and game feel. Other users said one large prompt could burn a 5-hour window and 20% of a weekly $20 plan.

NEWS3w ago
Kimi K3 benchmarks question cost advantage as token use rises

Illscience found Kimi K3 strong on prose reasoning but costly per task because it used more tokens than expected. LLMJunky's Rocket League-style test said the model made good UI but lacked Fable-level polish.

NEWS3w ago
Terra 5.6 high cuts Clawsweeper GitHub review time by ~40%, Steipete says

Steipete said he moved the Clawsweeper GitHub review bot to 5.6 Terra high and saw about 40% faster reviews at lower cost with little quality change. A follow-up framed the result as a warning that general model benchmarks may not predict issue-review performance.

RELEASE3w ago
Kimi K3 gets creator tests against GPT-5.6 Sol and Fable 5

Creators posted open-weights benchmarks and tests comparing Kimi K3 with GPT-5.6 Sol and Fable 5. Demos covered UI animation, games, Seedance renders, kernel code, and reported OpenRouter rate limits.

RELEASE4w ago
BytePlus previews Seedance 2.5 with Michael Owen football footage

BytePlus posted a Michael Owen football recreation made with Seedance 2.5 as creators treated it as a realism benchmark. Posts cite a Friday launch target and 30-second generations, pending no delay.

NEWS4w ago
Goodside benchmarks GPT-5.6 Sol Pro on 150- and 1,025-Pokémon crossword tests

Goodside compared Claude Fable 5 Max puzzle generation with GPT-5.6 Sol Pro solving attempts. Sol solved a 150-Pokémon empty crossword but failed the 1,025-Pokémon version, with one reported success traced to the answer key.

NEWS4w ago
Curious Refuge ranks GPT Image 2 above Meta Muse in character tests

Curious Refuge compared Meta’s Muse image model with GPT Image 2 across character sheets, infographics, posters, cinematic stills, edits, and book covers. It favored GPT Image 2 for realism and consistency.

NEWS4w ago
Goodside tests GPT-5.6 Sol on fake handwriting and Ghost Font prompts

Goodside tested GPT-5.6 Sol and Claude Fable 5 with fake handwriting, constrained-vocabulary prompts, and Ghost Font. Sol often answered unreadable inputs, while Fable more often refused or pushed back.

NEWS4w ago
Goodside benchmarks GPT-5.6 Sol hallucinations on binary noise

Goodside fed GPT-5.6 Sol and Claude Fable 5 binary noise and meaningless handwriting. Sol often invented hidden text, while Fable also failed on noise but more often pushed back on scribbles.

RELEASE4w ago
GPT-5.6 Sol gets creator tests in Figma Make and Fable 5 comparisons

GPT-5.6 Sol appeared in Figma Make and creator benchmarks against Fable 5 across design, games, and video. Testers praised browser persistence and design output, while noting higher token use and mode confusion.

WORKFLOW1mo ago
Kaigani builds video-QC benchmark after Gemini and Codex frame tests

Kaigani posted generated-video QC tests where Gemini caught subtle errors and said they are deliberately glitching clips for a benchmark. The work follows earlier Codex frame-analysis tests and focuses on detecting defects in generated video.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.