SpecFold speeds up diffusion LLM speculative decoding by up to 1.99x

Diffusion large language models (DLLMs) generate text by iterative block denoising. Multi-branch speculative decoding speeds this up by verifying a main branch together with several draft branches in a single forward pass. A new paper argues that this verification step wastes work, and proposes SpecFold to fix it.
The authors note that earlier DLLM acceleration methods mostly exploit temporal redundancy, meaning similarity across denoising steps. They identify a second, complementary axis inside each speculative verification step, which they call multi-branch computational redundancy. Draft branches inherit most tokens from their parents and unmask only a small set of additional positions, so large portions of the hidden states stay highly similar from branch to branch.
SpecFold is described as an algorithm-system co-design that exploits this overlap. On the algorithm side, it performs token-level residual gating and selectively reuses parent computation through folded attention and FFN, while preserving residual hidden states. On the systems side, a Triton kernel implementation turns this fine-grained reuse into end-to-end throughput gains through efficient sparse multi-branch execution.
The authors say SpecFold is orthogonal to temporal caching and compatible with existing DLLM speculation strategies, so it is meant to stack with those approaches rather than replace them.
Across two DLLM families, five models and five standard benchmarks, SpecFold achieves up to 1.64x throughput over Spiffy and up to 1.99x over vanilla decoding, while maintaining comparable task performance.
Key facts
- SpecFold targets multi-branch computational redundancy: draft branches inherit most tokens from their parents, so hidden states stay highly similar across branches.
- It combines token-level residual gating with folded attention and FFN that selectively reuse parent computation, plus a Triton kernel for sparse multi-branch execution.
- Reported speedups are up to 1.64x throughput over Spiffy and up to 1.99x over vanilla decoding.
- Evaluation covers two DLLM families, five models and five standard benchmarks, with task performance described as comparable.
- The authors say the method is orthogonal to temporal caching and compatible with existing DLLM speculation strategies.
Why it matters
Speculative decoding is one of the main ways to make diffusion language models faster, and most prior acceleration work has focused on similarity across denoising steps. This paper points at a different source of waste: the near-duplicate computation among draft branches inside a single verification pass. Removing it attacks the cost of verification itself, and the authors say it can be combined with temporal caching.
Who it affects
The work is aimed at people building and serving diffusion LLMs, especially those already using multi-branch speculative decoding. Because the authors describe SpecFold as compatible with existing DLLM speculation strategies, it is relevant to anyone running such strategies who wants more throughput from the same models.
How to use it
The abstract describes the technique as folded attention and FFN with token-level residual gating, backed by a Triton kernel for sparse multi-branch execution. No code release, license or publication venue is mentioned, so there is nothing to download or install from the source as given. Practitioners would need to wait for an implementation or reproduce the method from the paper.
How solid is it
The claims come from the paper's own abstract, and the evaluation is broad on paper: two DLLM families, five models and five standard benchmarks. The 1.64x and 1.99x figures are stated as 'up to' maxima, and no average or median speedup is given. The abstract gives no numeric task-performance scores, so 'comparable' is not quantified. It also does not name the DLLM families, the models or the benchmarks, and Spiffy is described only as the comparison baseline.
Risks and caveats
Treat the headline numbers as best-case results. The abstract gives no hardware, batch size, sequence length or absolute tokens-per-second figures, so real-world gains on a given setup are unknown. Preserving 'comparable' task performance is the authors' claim and cannot be checked from the abstract alone. The abstract names no authors or institutions, and the results have not been independently reproduced.
“SpecFold is orthogonal to temporal caching and compatible with existing DLLM speculation strategies.”
— SpecFold paper abstract