MODEL3 stories
DeepSeek V4 Flash
Stories, products, and related signals connected to this tag in Explore.
Stories
Filter storiesRELEASE3rd August
DeepSeek V4 Flash adds Baseten and Together AI serving with 1M-token context
Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.
NEWS2nd August
Cline raises free DeepSeek Flash quota 3x for coding agents
Cline tripled its free quota while Nous discounted Flash 0731 by 90% for a week and OpenHands offered free cloud use. OpenCode reported 8T Flash tokens on Aug. 1.
NEWS1st August
DeepSeek V4 Flash benchmarks show cheaper tokens but 3x SWE-Bench task cost
New tests showed DeepSeek V4 Flash as cheaper per token and faster on some serving paths. Ramp said it cost 3x more than GPT-5.6 Luna per SWE-Bench task because it used more turns.