• Sources: Level1Techs write-up, HN discussion
  • Summary: The author holds hardware, weights, drivers, and prompt fixed and changes only the vLLM attention backend, reporting bit-identical logits within a backend and reproducible top-1 token flips between backends that turn into wrong Cisco commands. An NVIDIA NVFP4 checkpoint reached roughly 50 percent token flips at 88k context and failed to close tool calls. The parts were published from 2026-08-16 through 2026-08-20 by a forum user and are not peer reviewed, with the methodology disclosed in enough detail to check.
  • Why it matters: Agent correctness depends on inference stack configuration and not only on the model, so a backend or quantization choice can produce reproducible wrong tool calls in production.

send feedback on this story