- Sources: vLLM blog, HN discussion
- Summary: The post on the vLLM blog, bylined AMD and Embedded LLM and dated 2026-08-23, reports measured speculative decoding throughput on AMD MI300X and MI355X under ROCm. It compares five drafting methods: native MTP, Gemma 4 MTP, EAGLE-3, DFlash, and DSpark. The post states the throughput effect varied by drafting method, proposal length, model family, draft checkpoint, workload, and acceptance behavior.
- Why it matters: Speculative decoding on ROCm is a per-workload tuning exercise across drafting method, proposal length, and draft checkpoint rather than a switch that is turned on.
- Follow-up: Whether the same drafting-method ranking holds on other model families and on non-ROCm hardware.
send feedback on this story