"Unlocking Lossless Speedups in LLMs via Discrete Diffusion" by Subham Sekhar Sahoo , Lingjie Chen , Khiem Pham , Jonathan Geuter , Chaitanya Dwivedi , Varad Pimpalkhute , Yash Akhauri , Alexander Moreno , Mikhail Yurochkin , Zhenting Wang , Mostafa Elhoushi , Nolan Dey , Shane Bergsma , Joel Hestness , John Thickstun , Eric Xing , Zhengzhong Liu
TLDR:
Diffusion-augmented autoregressive language models, such as Uno, improve upon Large Language Models (LLMs) by introducing a new approach that accelerates inference without compromising quality. These models use diffusion to draw multiple tokens in parallel from the autoregressive model distribution, decoupling AR weights and lightweight diffusion weights for more efficient training. By introducing the Ψ-Spec sampler, Uno achieves lossless acceleration and inference-time scaling without the need for a separate draft model or sacrificing AR model quality. Uno outperforms leading d-LLMs and proprietary models across various benchmarks, delivering up to 3 times speedups and higher throughput at different batch sizes. The code and checkpoints for Uno are available for public use.
Read Paper / Blog

huggingface.co
Paper page - Unlocking Lossless Speedups in LLMs via Discrete Diffusion
Join the discussion on this paper page
1September 8, 2026 36