DSpark
Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
DSpark is DeepSeek-AI's speculative decoding framework/drafter approach for accelerating LLM inference. It combines semi-autoregressive parallel draft generation with confidence-scheduled, load-aware verification, and DeepSeek released DSpark checkpoints and implementations through the DeepSpec repository.

Recent stories
A practitioner report found speculative-decoding acceptance fell from 0.71 to 0.18 beyond about 32K context. Separate DSpark and mlx-dspark tests reported speedups on RTX and Apple Silicon setups.
Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.