• Sources: primary, discussion
  • Summary: The argument is that the model controls the token sequence an inference engine parses, and that engines supporting more than 200 model architectures and about 35 chat templates carry parser bugs reachable from that sequence. The essay cites a prior vLLM tool-call parser that passed arguments to eval(), and a separate benign mis-parse of a reasoning tag, as evidence that engines do more than map tokens to strings. The proposed defence is architectural: have the GPU host emit only logits, and sample, parse and forward on a second host. The essay is analysis rather than a disclosed vulnerability, and the CVE identifier it cites for the vLLM eval() bug did not resolve in NVD when checked for this run, so that identifier is carried as the essay's claim and not as an established record.
  • Why it matters: The split-host proposal is a concrete deployment decision for anyone self-hosting open-weight models behind vLLM or SGLang.

send feedback on this story