A transformer writes one token at a time, and every token needs a full pass through every layer. That is why a giant model feels slow: you are not waiting on thinking, you are waiting on the same enormous pile of weights being hauled out of memory again and again.
🐇 Speculative decoding cheats the queue. A small, cheap draft model guesses the next handful of tokens. The big model then checks all of them in one pass — verification is parallel even though generation is not. Guesses that match are kept, the first wrong one is corrected, and the draft starts running again.
The original paper by Yaniv Leviathan, Matan Kalman and Yossi Matias, published in November 2022, reported a 2x-3x speedup on T5-XXL with identical outputs — no retraining, no architecture change, no new model.
That last part is the whole point. Most speed tricks cost you something: quantisation, pruning, a smaller model, a shorter answer. This one produces mathematically the same text, just sooner.
🎯 Which is why a fair few "our model got faster this month" announcements really mean "we finally shipped a better guesser" 🙃
