Skip to content

DeepSeek Releases DSpark to Speed LLM Inference 60-85% with Speculative Decoding

DeepSeek released DSpark, an open-source speculative decoding framework that cuts DeepSeek-V4 per-user generation latency by 60-85% in production, alongside the DeepSpec training codebase.

DeepSeek whale logo in white on the official DeepSeek blue background
DeepSeek logo. Composite: techshooked.

DeepSeek released DSpark, a speculative decoding framework for large-model inference, alongside open-source checkpoints and the DeepSpec training codebase under an MIT license. In production on DeepSeek-V4, the framework reduces per-user generation latency by 60-85% compared to the prior single-token MTP-1 baseline, without changing the output distribution or retraining the target model.

Speculative decoding splits generation between a small draft model, which proposes a block of tokens cheaply, and the full target model, which verifies the block in one forward pass. The key tension is acceptance rate: parallel drafters produce tokens quickly but see rapid quality decay as the block lengthens; autoregressive drafters maintain high acceptance but grow more expensive with block size. DSpark addresses this by splitting the draft into two stages. A parallel backbone generates base logits across all positions simultaneously; a lightweight sequential head then applies a prefix-dependent bias before sampling each token, using the immediately preceding token to stabilize acceptance deep into the block. A low-rank factorization keeps the sequential head cheap even at large vocabulary sizes.

A second component, a confidence-aware verification scheduler, adjusts how many draft tokens are verified per request based on measured GPU load. When the system is underutilized, it verifies longer prefixes; when concurrency is high, it trims the budget to protect throughput. This is where the production gains are largest: on Flash and Pro variants of DeepSeek-V4, per-user generation runs 60-85% faster than MTP-1 at matched throughput, according to DeepSeek's own production measurements.

The DeepSpec repository ships with three draft-model implementations: DSpark, DFlash, and Eagle3, along with data preparation utilities, training code, and a benchmark evaluation suite covering nine datasets including GSM8K, HumanEval, and AlpacaEval. The default configuration assumes a single node with 8 GPUs. A notable infrastructure caveat: caching target model outputs for training can require roughly 38 TB of storage for the default Qwen3-4B setting.

Production checkpoints for DeepSeek-V4-Pro-DSpark and DeepSeek-V4-Flash-DSpark attach a draft module to existing V4 weights, meaning teams already running DeepSeek-V4 can adopt DSpark without changing their base model. The shipped configuration uses five-token draft blocks with the Markov head. The broader question is whether DSpark's two-stage drafting strategy generalizes well to models outside the DeepSeek family; the paper benchmarks it against Qwen3 and Gemma targets, which offers an early indication but does not cover the full range of production deployments.

Share this story

Julian Beaumont

Julian Beaumont covers artificial intelligence and large language models for techshooked, following the path from research paper to deployed feature. His standard is anti-hype: ask what a model actually does, what trained it, how it fails, and whether a benchmark measures what the announcement claims.