- Sources: arXiv 2607.24555
- Summary: The paper proposes attaching a spectral summary to each page of the KV cache and selecting pages from that resident index, which it puts at about a tenth of the cache size, so selection reads no candidate keys or values. It reports decode latency at a 1M-token context cut to about half. It is packaged as a drop-in plugin for unmodified vLLM, with batched decode running in full CUDA graphs. Every figure here is the paper's own. This is a first-version preprint with no independent reproduction and no discussion thread.
- Why it matters: Long-context serving cost is dominated by reading the whole KV cache at every decode step, and a selection path that ships against an unmodified serving stack is testable by a team rather than only by the authors.
- Follow-up: Watch for an independent run of the plugin against a stock vLLM deployment.
send feedback on this story