• Sources: primary, discussion
  • Summary: The release page describes Bonsai 2 27B as a ternary compression of Qwen3.8 27B, distributed at 5.9GB under Apache 2.0, retaining 98.2% of aggregate benchmark performance against 95% retention in the prior generation. The same page states that the model runs on NVIDIA GPUs through CUDA and on Apple devices through MLX, in both cases through custom low-bit kernels. Every capability claim is the publisher's own and none is independently reproduced here.
  • Why it matters: Near-parity benchmarks on a 27B-class model in 5.9GB put coding-agent and computer-use workloads on a single consumer GPU or on Apple silicon.
  • Follow-up: Track independent benchmark results reproducing the 98.2% retention claim.

send feedback on this story