• Sources: primary, discussion
  • Summary: Cloudflare published per-concurrency measurements rather than a single headline number. An FP8 KV cache loses a few percent per token but admits 64 concurrent requests where BF16 exhausts memory at 32. INT4 weights shrink a GLM 5.2 checkpoint from 705 GB to 421 GB with accuracy within 0.8 points.
  • Why it matters: The numbers state the quantization tradeoff as a concurrency ceiling rather than a quality score, which is the form the decision actually takes when sizing inference capacity.

send feedback on this story