- Sources: Inco blog, HN discussion, r/LocalLLaMA throughput report, two RTX 3090, r/LocalLLaMA throughput report, single RTX 3090
- Summary: Inco reports 2.7 to 3.4 times the throughput of autoregressive decoding on Qwen3.8-27B at batch size 1, for roughly 1.3 percent added draft-verify cycle latency. The speedup is lossless, because rejection sampling restores the exact target distribution, so this changes throughput and not output. The path selector adds 2.0 million parameters against 77.8 million for the DSpark correction it is compared to, and the post states that the comparison drafters were trained by Inco itself. It already runs in SGLang, vLLM, llama.cpp, Ollama, and oMLX, and independent r/LocalLLaMA posts report 134 tokens per second on a single RTX 3090 and 218 on two.
- Why it matters: A lossless decoding change already merged into five serving stacks raises throughput without a model swap or an output-quality tradeoff.
send feedback on this story