- Sources: primary, discussion
- Summary: A benchmark of Qwen3.8 27B across quantization levels, posted 2026-08-26 and picked up on Hacker News on 2026-09-08, reports 4-bit output at parity with BF16 on Terminal-Bench 2.1 and GPQA Diamond, and 1-bit output at around random chance on GPQA Diamond. The three benchmarks are GPQA Diamond, IFBench and Terminal-Bench 2.1, and the run replicated the official BF16 numbers first, so the quantized deltas are measured against a reproduced baseline. The post puts BF16 at 55 GB and Q4_K_M at 17 GB, which fits a 24 GB card such as an RTX 4090 with room for about 64k tokens of context.
- Why it matters: It puts a number on what a local coding model costs in GPU memory, 17 GB at Q4_K_M against 55 GB at BF16, which is the difference between one 24 GB card and a multi-GPU host.
send feedback on this story