DeepSeek-V4-Flash-0731
The official release of DeepSeek-V4-Flash with enhanced agentic capabilities.
A specific language-model release identified in the supplied evidence as DeepSeek V4 Flash / Flash-0731, used for coding-agent and long-context workloads.
Pricing
Peak rate (01:00–04:00 and 06:00–10:00 UTC): $0.44/M cache-miss input, $1.32/M output, and $0.014/M cache-hit input. Off-peak rate: $0.22/M cache-miss input, $0.66/M output, and $0.007/M cache-hit input. Official page says these rates took effect August 16, 2026.
DeepSeek’s official API pricing table identifies the deepseek-v4-flash model version as DeepSeek-V4-Flash-0731. As of the stated effective date (August 16, 2026), it uses peak/off-peak token billing; normalized fields record peak prices.
Model Intelligence
Recent stories
DeepSeek V4.1 Flash leads Vals and Artificial Analysis open-weight comparisons, according to the evaluators. Its encoder-decoder design shares compressed KV state across decoder layers to reduce serving costs.
A practitioner report found speculative-decoding acceptance fell from 0.71 to 0.18 beyond about 32K context. Separate DSpark and mlx-dspark tests reported speedups on RTX and Apple Silicon setups.
OpenCode’s thdxr said Go users spent $1.14 per day on DeepSeek V4 Flash last week. Wafer added a fast OpenRouter route, while Nous extended a 90% discount for the 0731 model.
Together says two V4 Flash attempts solved more DeepSWE tasks than one GPT-5.6 Luna attempt for roughly one-third the cost. Practitioners report Flash-0731 results vary sharply by harness and pass count.
Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.
Cline tripled its free quota while Nous discounted Flash 0731 by 90% for a week and OpenHands offered free cloud use. OpenCode reported 8T Flash tokens on Aug. 1.
New tests showed DeepSeek V4 Flash as cheaper per token and faster on some serving paths. Ramp said it cost 3x more than GPT-5.6 Luna per SWE-Bench task because it used more turns.
DeepSeek released V4 Flash 0731 with weights, a technical report, API access, 1M context, MoE routing, and low token prices. Its cited benchmarks show gains on Artificial Analysis, Terminal-Bench, Frontend Code Arena, and agent tests.