- Sources: PrismML write-up, HN 48910545
- Summary: PrismML published Bonsai 27B, an extreme-quantization build of Qwen 3.6 27B that reduces most weights to ternary values with group-wise FP16 scales, reaching about 1.71 effective bits per weight and shrinking the model from roughly 54 GB in FP16 to about 3.8 GB. The stated goal is on-device inference on high-memory phones such as recent iPhone Pro models, with reported speeds near 1 token per second on consumer hardware. PrismML reports roughly 90% capability retention versus the full model, stronger on math and code than a smaller Gemma build and weaker on knowledge, tool calling, and vision. Weights are posted publicly.
- Comments: HN commenters report the model gets stuck in reasoning loops and cite an independent perplexity measurement well above the unquantized baseline, and question whether the packing format used is the most efficient ternary representation. The retention and speed figures are vendor claims and are not independently reproduced.
- Why it matters: Sub-2-bit quantization that fits a 27B-class model in under 4 GB pushes on-device inference further, but the reported loop behavior and perplexity gap show the accuracy cost is not yet settled.
- Follow-up: Watch for independent perplexity and task benchmarks, the exact packing format, and reproduction of the on-device speed claim.
send feedback on this story