Benchmark
Stories, products, and related signals connected to this tag in Explore.
Stories
Filter storiesAlibaba released Qwen3.8-Max, described in a launch thread as a 2.4T sparse MoE with 95B active parameters, 1M context, native vision/text, and agent benchmarks. API pricing is listed at $2/$6 per 1M tokens, with open weights planned for Hugging Face.
Reddit tests show MiniMax H3 running locally on GPUs from RTX 3060 to 5090. A 3060 reportedly takes about an hour for 10 seconds, while a 5090 makes 15 seconds in 84 seconds, with users also reporting motion and reference-video failures.
Graft scans a repo into linked markdown and plugs into Claude Code hooks so the agent can reuse codebase context. Its 162-run benchmark claims 46% fewer tool calls, up to 4x fewer tokens, and 60% less time.
TopviewAI posts cite native 2K, 15-second generations, multimodal control, and claimed pricing at 30% of Seedance 2.0. Creators also tested H3 on ComfyUI Cloud, with ComfyUI support for open weights reported.
Flux 3 appeared in early preview on Nous Hermes Agent as an open-weights video model, with 720p sample tests circulating. Curious Refuge ranked its early results behind Seedance 2.0 and LTX 2.3, making output quality the main launch caveat.
Anthropic launched Claude Opus 5 for Claude Code and the API with high-effort defaults, fast mode, migration tooling, mid-conversation tool changes, and routing fallback. The model scored 43.3% on Frontier-Bench v0.1.
Google says Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber target faster, cheaper agent workloads. Early tests cite lower token use and stronger coding results.
Posts compared Kimi K3 with Fable on Rocket League-style app builds, finding stronger UI than physics and game feel. Other users said one large prompt could burn a 5-hour window and 20% of a weekly $20 plan.
Illscience found Kimi K3 strong on prose reasoning but costly per task because it used more tokens than expected. LLMJunky's Rocket League-style test said the model made good UI but lacked Fable-level polish.
Steipete said he moved the Clawsweeper GitHub review bot to 5.6 Terra high and saw about 40% faster reviews at lower cost with little quality change. A follow-up framed the result as a warning that general model benchmarks may not predict issue-review performance.
Creators posted open-weights benchmarks and tests comparing Kimi K3 with GPT-5.6 Sol and Fable 5. Demos covered UI animation, games, Seedance renders, kernel code, and reported OpenRouter rate limits.
BytePlus posted a Michael Owen football recreation made with Seedance 2.5 as creators treated it as a realism benchmark. Posts cite a Friday launch target and 30-second generations, pending no delay.
Goodside compared Claude Fable 5 Max puzzle generation with GPT-5.6 Sol Pro solving attempts. Sol solved a 150-Pokémon empty crossword but failed the 1,025-Pokémon version, with one reported success traced to the answer key.
Curious Refuge compared Meta’s Muse image model with GPT Image 2 across character sheets, infographics, posters, cinematic stills, edits, and book covers. It favored GPT Image 2 for realism and consistency.
Goodside tested GPT-5.6 Sol and Claude Fable 5 with fake handwriting, constrained-vocabulary prompts, and Ghost Font. Sol often answered unreadable inputs, while Fable more often refused or pushed back.
Goodside fed GPT-5.6 Sol and Claude Fable 5 binary noise and meaningless handwriting. Sol often invented hidden text, while Fable also failed on noise but more often pushed back on scribbles.
GPT-5.6 Sol appeared in Figma Make and creator benchmarks against Fable 5 across design, games, and video. Testers praised browser persistence and design output, while noting higher token use and mode confusion.
Kaigani posted generated-video QC tests where Gemini caught subtle errors and said they are deliberately glitching clips for a benchmark. The work follows earlier Codex frame-analysis tests and focuses on detecting defects in generated video.