- Sources: primary, discussion
- Summary: Cloudflare published per-concurrency measurements rather than a single headline number. An FP8 KV cache loses a few percent per token but admits 64 concurrent requests where BF16 exhausts memory at 32. INT4 weights shrink a GLM 5.2 checkpoint from 705 GB to 421 GB with accuracy within 0.8 points.
- Why it matters: The numbers state the quantization tradeoff as a concurrency ceiling rather than a quality score, which is the form the decision actually takes when sizing inference capacity.
send feedback on this story