Skip to content
AI Primer

DeepSeek-V4-Flash-0731

Official DeepSeek-V4-Flash release with enhanced agentic capabilities and 1M-token context.

DeepSeek-V4-Flash-0731 is the official 2026-07-31 release of DeepSeek-V4-Flash, superseding the preview version with enhanced agentic capabilities while keeping the same architecture and size as the preview. It is served through DeepSeek's API as model ID deepseek-v4-flash and published as open weights on Hugging Face.

Pricing

Official site · Aug 9, 2026, 7:11 AM
Input / 1M
$0.14
Output / 1M
$0.28
Cached input / 1M
$0.0028

Official page says prices are per 1M tokens and billing is based on total input/output tokens. Input price recorded is cache-miss price; cached input price recorded separately. DeepSeek notes pricing may increase in the near future and recommends checking the page for current pricing.

DeepSeek's official API pricing page lists model version DeepSeek-V4-Flash-0731 under model deepseek-v4-flash, with prices explicitly stated per 1M tokens: cache-hit input $0.0028, cache-miss input $0.14, and output $0.28.

View source

Model Intelligence

Context window
1,000,000 tokens
Arena ranking
52
Benchmarkable
Yes
Model level
release
Intelligence Index
51.8
Coding Index
69.1
GPQA
0.91
HLE
0.39
SciCode
0.5
LCR
0.74

Recent stories

6 linked stories
newsPRIMARY2026-08-09
OpenCode users reportedly average $1.14/day on DeepSeek V4 Flash

OpenCode’s thdxr said Go users spent $1.14 per day on DeepSeek V4 Flash last week. Wafer added a fast OpenRouter route, while Nous extended a 90% discount for the 0731 model.

newsPRIMARY2026-08-08
DeepSeek V4 Flash benchmarks claim lower DeepSWE cost than GPT-5.6 Luna

Together says two V4 Flash attempts solved more DeepSWE tasks than one GPT-5.6 Luna attempt for roughly one-third the cost. Practitioners report Flash-0731 results vary sharply by harness and pass count.

releaseSECONDARY2026-08-03
DeepSeek V4 Flash adds Baseten and Together AI serving with 1M-token context

Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.

newsSECONDARY2026-08-02
Cline raises free DeepSeek Flash quota 3x for coding agents

Cline tripled its free quota while Nous discounted Flash 0731 by 90% for a week and OpenHands offered free cloud use. OpenCode reported 8T Flash tokens on Aug. 1.

newsSECONDARY2026-08-01
DeepSeek V4 Flash benchmarks show cheaper tokens but 3x SWE-Bench task cost

New tests showed DeepSeek V4 Flash as cheaper per token and faster on some serving paths. Ramp said it cost 3x more than GPT-5.6 Luna per SWE-Bench task because it used more turns.

releasePRIMARY2026-07-31
DeepSeek releases V4 Flash 0731 as MIT-licensed open weights

DeepSeek released V4 Flash 0731 with weights, a technical report, API access, 1M context, MoE routing, and low token prices. Its cited benchmarks show gains on Artificial Analysis, Terminal-Bench, Frontend Code Arena, and agent tests.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.