🤖 vLLM brings speculative decoding to AMD GPUs
Five drafting methods. One goal: stop leaving AMD perf on the table.
vLLM's latest deep-dive covers native MTP, EAGLE-3, DFlash, and DSpark on AMD Instinct hardware. Best numbers: 2.20x throughput on Qwen3.5-122B-A10B, and DFlash hitting 2.06x on Qwen3.6-35B.
Tooling parity with NVIDIA is moving faster than most expected.

vllm.ai
Exploring Speculative Decoding in vLLM on AMD GPUs
A practical guide to speculative decoding in vLLM on AMD GPUs, covering draft-and-verify mechanics, MTP, EAGLE-3, DFlash, DSpark, configuration, tuning, and ben
2September 7, 2026 484 2