• Sources: neomindlabs.com, ik_llama.cpp PR 2138, HN discussion
  • Summary: A write-up dated 2026-06-08, which reached the HN front page on 2026-07-15, runs Google's Gemma 4 26B-A4B Mixture-of-Experts model quantized to Q8_0 on a dual Xeon E5-2690 v2 (Ivy Bridge, 2013, AVX1 only, no AVX2 or FMA3, no GPU) at about 5.2 tokens per second decode. The build produced fluent-looking multilingual gibberish because two fused MoE graph ops (MOE_FUSED_UP_GATE, FUSED_UP_GATE) were still emitted by the graph builder but had no dispatch case on non-AVX2 builds, leaving roughly 240 tensors per forward pass uninitialized with mean logits pinned near +16. The fix splits the fused ops into ggml_mul_mat_id plus ggml_fused_mul_unary calls that have non-optimized implementations.
  • Comments: The author explained the diagnosis in the thread and posted the upstream fix as ik_llama.cpp PR 2138. Another commenter reported 8 to 12 tokens per second on comparable old hardware.
  • Why it matters: The bug is a concrete case of a code path silently skipped on an unusual build target, and it documents that large MoE models can run at reading speed on decade-old CPUs.

send feedback on this story