- Sources: arXiv 2607.21535
- Summary: A preprint applies a StreamingLLM sliding window with an attention sink to the attention of the speculative decoding draft head alone, leaving the verification pass at full attention. The authors report per-decode-step cost falling 28 percent to 44 percent at one million tokens of context, measured on three model families in SGLang. The method is training-free. The authors argue it is lossless because the full-attention verification path is untouched, so the distribution of accepted tokens does not change. The figures are the authors' own and are not independently reproduced.
- Why it matters: The draft model's KV cache is a cost most long-context serving stacks pay without measuring it separately, and the claim here is that it can be windowed without moving output quality.
send feedback on this story