- Sources: primary, discussion
- Summary: Holding the whole query plan up front turns KV cache eviction and join ordering into decisions the engine can make rather than guess, because it knows when an entry is dead and what demand follows. The AI.IF workload needs only the single token the final prefill pass emits, so the engine implements no decode phase, no sampling, no CUDA graph capture and no speculative decoding. The two headline figures are not interchangeable: over 1 billion tokens per minute per H100 and more than 10x over the vLLM baseline come from one multi-join query the authors state is where planning matters most, while 1.84x is the geometric mean over their released benchmark, which includes two queries they designed to expose remaining weaknesses.
- Why it matters: Code ships as quail-engine 0.1.0 with the forward passes forked from vLLM and the benchmark released, so the smaller figure is the one a team can check.
send feedback on this story