Skip to content
AI Primer

A specific language-model release identified in the supplied evidence as DeepSeek V4 Flash / Flash-0731, used for coding-agent and long-context workloads.

Pricing

Official site · Sep 6, 2026, 7:17 AM
Input / 1M
$0.44
Output / 1M
$1.32
Cached input / 1M
$0.014

Peak rate (01:00–04:00 and 06:00–10:00 UTC): $0.44/M cache-miss input, $1.32/M output, and $0.014/M cache-hit input. Off-peak rate: $0.22/M cache-miss input, $0.66/M output, and $0.007/M cache-hit input. Official page says these rates took effect August 16, 2026.

DeepSeek’s official API pricing table identifies the deepseek-v4-flash model version as DeepSeek-V4-Flash-0731. As of the stated effective date (August 16, 2026), it uses peak/off-peak token billing; normalized fields record peak prices.

View source

Model Intelligence

Context window
1,000,000 tokens
Arena ranking
35
Benchmarkable
Yes
Model level
release
Intelligence Index
34.5
Coding Index
69.1
GPQA
0.91
HLE
0.39
SciCode
0.5
LCR
0.8

Recent stories

8 linked stories
releaseSECONDARY2026-09-10
DeepSeek V4.1 Flash tops independent open-weight evaluations

DeepSeek V4.1 Flash leads Vals and Artificial Analysis open-weight comparisons, according to the evaluators. Its encoder-decoder design shares compressed KV state across decoder layers to reduce serving costs.

workflowSECONDARY2026-08-12
Speculative decoding tests report acceptance drop from 0.71 to 0.18 after ~32K context

A practitioner report found speculative-decoding acceptance fell from 0.71 to 0.18 beyond about 32K context. Separate DSpark and mlx-dspark tests reported speedups on RTX and Apple Silicon setups.

newsPRIMARY2026-08-09
OpenCode users reportedly average $1.14/day on DeepSeek V4 Flash

OpenCode’s thdxr said Go users spent $1.14 per day on DeepSeek V4 Flash last week. Wafer added a fast OpenRouter route, while Nous extended a 90% discount for the 0731 model.

newsPRIMARY2026-08-08
DeepSeek V4 Flash benchmarks claim lower DeepSWE cost than GPT-5.6 Luna

Together says two V4 Flash attempts solved more DeepSWE tasks than one GPT-5.6 Luna attempt for roughly one-third the cost. Practitioners report Flash-0731 results vary sharply by harness and pass count.

releaseSECONDARY2026-08-03
DeepSeek V4 Flash adds Baseten and Together AI serving with 1M-token context

Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.

newsSECONDARY2026-08-02
Cline raises free DeepSeek Flash quota 3x for coding agents

Cline tripled its free quota while Nous discounted Flash 0731 by 90% for a week and OpenHands offered free cloud use. OpenCode reported 8T Flash tokens on Aug. 1.

newsSECONDARY2026-08-01
DeepSeek V4 Flash benchmarks show cheaper tokens but 3x SWE-Bench task cost

New tests showed DeepSeek V4 Flash as cheaper per token and faster on some serving paths. Ramp said it cost 3x more than GPT-5.6 Luna per SWE-Bench task because it used more turns.

releasePRIMARY2026-07-31
DeepSeek releases V4 Flash 0731 as MIT-licensed open weights

DeepSeek released V4 Flash 0731 with weights, a technical report, API access, 1M context, MoE routing, and low token prices. Its cited benchmarks show gains on Artificial Analysis, Terminal-Bench, Frontend Code Arena, and agent tests.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.