DeepSeek released DSpark, a speculative decoding framework for large-model inference, alongside open-source checkpoints and the DeepSpec training codebase under an MIT license. In production on DeepSeek-V4, the framework reduces per-user generation latency by 60-85% compared to the prior single-token MTP-1 baseline, without changing the output distribution or retraining the target model.
Speculative decoding splits generation between a small draft model, which proposes a block of tokens cheaply, and the full target model, which verifies the block in one forward pass. The key tension is acceptance rate: parallel drafters produce tokens quickly but see rapid quality decay as the block lengthens; autoregressive drafters maintain high acceptance but grow more expensive with block size. DSpark addresses this by splitting the draft into two stages. A parallel backbone generates base logits across all positions simultaneously; a lightweight sequential head then applies a prefix-dependent bias before sampling each token, using the immediately preceding token to stabilize acceptance deep into the block. A low-rank factorization keeps the sequential head cheap even at large vocabulary sizes.
A second component, a confidence-aware verification scheduler, adjusts how many draft tokens are verified per request based on measured GPU load. When the system is underutilized, it verifies longer prefixes; when concurrency is high, it trims the budget to protect throughput. This is where the production gains are largest: on Flash and Pro variants of DeepSeek-V4, per-user generation runs 60-85% faster than MTP-1 at matched throughput, according to DeepSeek's own production measurements.
The DeepSpec repository ships with three draft-model implementations: DSpark, DFlash, and Eagle3, along with data preparation utilities, training code, and a benchmark evaluation suite covering nine datasets including GSM8K, HumanEval, and AlpacaEval. The default configuration assumes a single node with 8 GPUs. A notable infrastructure caveat: caching target model outputs for training can require roughly 38 TB of storage for the default Qwen3-4B setting.
Production checkpoints for DeepSeek-V4-Pro-DSpark and DeepSeek-V4-Flash-DSpark attach a draft module to existing V4 weights, meaning teams already running DeepSeek-V4 can adopt DSpark without changing their base model. The shipped configuration uses five-token draft blocks with the Markov head. The broader question is whether DSpark's two-stage drafting strategy generalizes well to models outside the DeepSeek family; the paper benchmarks it against Qwen3 and Gemma targets, which offers an early indication but does not cover the full range of production deployments.













