Skip to content
AI Primer
update

DeepSeek V4 Flash benchmarks at 61.4% on ARC-AGI-2 for $0.04 per task

ARC Prize verified DeepSeek V4 Flash at 61.4% on ARC-AGI-2 for $0.04 per task. Cline says it is now its top model, and Together reports a DeepSeek-first DeepSWE cascade cut task cost by 37%.

9 min read
DeepSeek V4 Flash benchmarks at 61.4% on ARC-AGI-2 for $0.04 per task
DeepSeek V4 Flash benchmarks at 61.4% on ARC-AGI-2 for $0.04 per task

TL;DR

  • ARC Prize verified DeepSeek V4 Flash 0731 at 89.0% on ARC-AGI-1 and 61.4% on ARC-AGI-2 at max effort, with task costs of $0.02 and $0.04, according to ARC Prize's result post.
  • The 0731 update is a re-post-trained Flash release, not a new API migration: Together's benchmark screenshot highlights Terminal-Bench 2.1 at 82.7 for Flash 0731 versus 72.1 for V4 Pro Preview.
  • Cheap rollouts change the coding benchmark math: Together's DeepSWE cascade test says a DeepSeek-first cascade solved more tasks than GPT-5.6 Luna alone at 37% lower cost per task.
  • Adoption moved fast inside coding tools: Cline's usage chart says DeepSeek V4-Flash became Cline's most-used model, and Ollama's rollout says 0731 is the new cloud default with 120+ output tokens per second.
  • The current price floor already has an asterisk: TestingCatalog's notice screenshot shows DeepSeek warning that overall API pricing will rise significantly.

DeepSeek's changelog says the API name stayed deepseek-v4-flash, the architecture stayed the same, and the model was only re-post-trained. ARC Prize's full result page publishes the low, high, and max reasoning variants, plus pass/fail rows for all 120 ARC-AGI-2 public tasks. DeepSeek's pricing page now carries the significant-increase footnote, while its context caching docs explain the disk cache that makes repeated prefixes unusually cheap.

What changed in 0731

DeepSeek's own framing is unusually specific: the public beta upgraded Flash only, kept the model name, left V4 Pro and app/web models unchanged, and added native Responses API support for Codex in the official V4 Flash API.

The concrete changes:

  • deepseek-v4-flash now routes to DeepSeek-V4-Flash-0731, according to DeepSeek's changelog.
  • The model keeps the same architecture and size as the Flash preview, and was only re-post-trained, according to the same changelog.
  • DeepSeek's Hugging Face model card calls 0731 the official release superseding the preview, with MIT-licensed weights.
  • The model card lists low, high, and max reasoning effort levels.
  • The model card says DSpark speculative decoding can be enabled in vLLM with --speculative-config.
  • DeepSeek recommends temperature = 1.0, top_p = 0.95 for agentic scenarios, and up to 384K max output for high and max local runs.

The benchmark table on the model card puts Flash 0731 above Flash Preview and V4 Pro Preview on Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, Toolathlon-Verified, Agents' Last Exam, AutomationBench Public, DSBench-FullStack, and DSBench-Hard.

ARC-AGI scores

ARC Prize's verified result is the cleanest headline number because it reports both performance and dollars per task.

ARC Prize's result page lists three reasoning variants:

  • Max: ARC-AGI-1 89.0%, ARC-AGI-2 61.4%, ARC-AGI-3 not listed.
  • High: ARC-AGI-1 87.0%, ARC-AGI-2 56.0%, ARC-AGI-3 not listed.
  • Low: ARC-AGI-1 84.0%, ARC-AGI-2 46.0%, ARC-AGI-3 not listed.

ARC Prize said ARC-AGI-3 evaluations are more operationally intensive and will roll out over the next few weeks in its follow-up. Its testing policy says public testing results include model outputs, durations, costs, and per-task scores, while frontier-model tests use a Semi-Private Evaluation Set with zero-data-retention agreements.

The oddity in the thread is that max reasoning was cheaper than high on at least one ARC task. ARC Prize's task note said high spent 98K tokens pursuing the wrong basin interpretation on task 28a6681f, while max spent 76K tokens and solved it.

DeepSWE cascade

Together's DeepSWE run is the useful engineering benchmark because it measures rollouts on 113 real long-horizon coding tasks from live open-source repos, four trials each, with hidden test-suite grading.

The full Together analysis reports:

  • Pass@1: GPT-5.6 Luna 67.2%, DeepSeek V4 Flash 53.3%.
  • Cost per rollout: Luna $0.61, DeepSeek $0.10.
  • Solves per $100: Luna 110, DeepSeek 532.
  • Failure regressions: Luna broke existing tests in 15% of failures, DeepSeek in 9%.
  • Median runtime: Luna 16 minutes, DeepSeek 23 minutes.
  • Median output: Luna 70K tokens, DeepSeek 104K tokens.
  • Cascade: DeepSeek first, Luna on failure, solved 78.9% at $0.385 per task.

Together's follow-up phrased the same result as DeepSeek delivering 80% of Luna's performance at roughly one-sixth the cost in its DeepSWE thread. That is the whole trick: the weaker model is cheap enough to be used as the first throw.

Harness choice

Composio's agent harness test is the reminder that model benchmarks are not agent benchmarks.

Composio ran DeepSeek V4 Flash through Claude Code, Codex, OpenCode, and Oh My Pi on 30 multi-step tasks over Gmail, Sheets, Airtable, GitHub, Slack, Calendar, Notion, and PagerDuty in its task description. The headline table split three ways:

  • Success rate: Oh My Pi 17/30, Claude Code 16/30, Codex 16/30, OpenCode 14/30.
  • Cost per successful task: OpenCode $0.073, Codex $0.081, Oh My Pi $0.103, Claude Code $0.195.
  • Median time: Claude Code 122.7s, OpenCode 129.7s, Codex 245.0s, Oh My Pi 272.4s.

Composio said seven tasks passed or failed depending only on the harness in its success-rate post. In a later reply, Composio described Claude Code as terse and fastest, Codex as methodical and wordy, OpenCode as brisk and cheap but more failure-prone, and Oh My Pi as slower with higher variance.

The omission people noticed was Cline. Composio's Cline reply asked what made it better than the tested harnesses, while another Composio reply acknowledged Cline is also popular.

Availability

DeepSeek V4 Flash 0731 showed up everywhere developers route coding traffic: Cline, Ollama, Pi, Hugging Face, Baseten, Vercel, CoreWeave, OpenRouter, and local GGUF stacks.

  • Cline said DeepSeek V4-Flash became its number one model, with usage up 40% since the 0731 update and tokens tripled in its chart thread.
  • Ollama said 0731 became the default for deepseek-v4-flash on Ollama Cloud, with zero-data-retention hosting in the US and Europe and 120+ output tokens per second in its rollout post.
  • Ollama also called it the fastest-growing model ever by token usage on Ollama and said it was scaling US and Europe capacity in an earlier capacity note.
  • Pi said DeepSeek V4 Flash was available across providers and called it Ollama's fastest-growing model plus the week's most popular model on OpenRouter in its adoption post.
  • Hugging Face's Baseten announcement says Baseten now serves the latest DeepSeek V4 Flash through Hugging Face Inference Providers with Python, JS, and agent-harness examples.
  • Baseten's post said developers can run Kimi K3, DeepSeek V4 Flash, and GLM-5.2 from Hugging Face model pages or with an HF token.
  • Baseten also pitched SFT, DPO, PPO, and GRPO for specializing DeepSeek V4 Flash on Baseten Loops in its training post.

Price hike and cache economics

DeepSeek still lists V4 Flash at $0.0028 per 1M cached input tokens, $0.14 per 1M cache-miss input tokens, and $0.28 per 1M output tokens, according to DeepSeek's pricing page. That same table now says overall API pricing will rise significantly, with the specific plan subject to official notice.

The pricing notice replaced an earlier peak-valley notice that would have doubled all billing items during peak hours in the before-and-after screenshot. Rohan Paul's pricing note tied the notice to the week after V4-Flash-0731 shipped and cited then-current rates of $0.14 and $0.28 per 1M input and output tokens.

DeepSeek's cache docs say context caching is enabled by default, builds a hard-disk cache for overlapping prefixes, reports prompt_cache_hit_tokens and prompt_cache_miss_tokens, and works on a best-effort basis. Cache construction takes seconds, and unused cache entries are usually cleared within a few hours to a few days, according to the official cache guide.

Developer-side traffic data made cache behavior part of the story. thdxr's cache table showed cache-hit ratios above 95% for ZCode, OpenCode V2, Cursor, Kimi Code CLI, Pi, Codex, Kilo Code, and OpenCode V1, while Claude Code / CLI sat at 89.31%.

Local runtime

Local V4 Flash is real, but the hardware bar is still home-lab weird.

Unsloth's DeepSeek V4 guide says Flash has 284B total parameters, 13B active parameters, and a 1M context window. It lists a 162GB lossless 8-bit GGUF, a 103GB 3-bit option that can run on a 110GB RAM device, and DSpark speedups for 0731.

Unsloth's tweet said DSpark made DeepSeek-V4-Flash-0731 GGUFs 1.4x to 2x faster with no accuracy change and could reach 120 tokens per second in its release post. Lars Grammel reported 35 to 40 generation tokens per second for a quantized 0731 build on a MacBook Pro M5.

Ben Davis' home setup was the other end of the spectrum. Davis' local Pi report described V4 Flash 0731 as usable for real work in a closet setup under $10,000, but still weaker than frontier models, slower, and limited to 2 to 4 sessions where his normal cloud workflow can involve 20+ Codex runs.

Vision gap

DeepSeek V4 Flash is text-only in the workflows people are using for coding agents, which showed up immediately in frontend work.

AI Builder Club called the price-to-performance "insane" but said the model cannot read a design mockup or screenshot-compare a frontend result without an added vision skill in its frontend-agent demo. In replies, AI Builder Club said vision remains essential for frontend work, and another reply said the same pattern can call other vision models too.

Yacine experimented with a multi-provider Codex setup where DeepSeek handled the main coding path and Luna supplied vision in his Codex post. He later said DeepSeek using Luna as eyes in Codex goal mode was surprisingly strong in a follow-up.

Agentic data in mid-training

The most interesting training breadcrumb is one sentence in the V4 paper: DeepSeek says it enhanced coding capabilities by incorporating agentic data during mid-training, inside a broader pretraining corpus of more than 32T tokens.

Chris Wolfe pointed out that the sentence is the only place the paper mentions the mid-training process used for DeepSeek V4, and read it as another sign that labs are blending post-training-style data into mid-training in his note. In replies, Wolfe said verified domains may allow rejection sampling of trajectories, while cases without deterministic verifiers get more complicated in the follow-up.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 9 threads
TL;DR2 posts
ARC-AGI scores2 posts
DeepSWE cascade1 post
Harness choice5 posts
Availability5 posts
Price hike and cache economics3 posts
Local runtime2 posts
Vision gap5 posts
Agentic data in mid-training1 post
Share on X