• Sources: primary, discussion
  • Summary: TurboFieldfare streams Gemma 4 mixture-of-experts weights from SSD instead of holding them resident. The project reports a 26B-parameter model running in about 2 GB of RAM. That figure is self-reported by the project and carries no external verification.
  • Why it matters: Streaming mixture-of-experts weights from SSD moves the local-inference constraint from RAM to storage and I/O.

send feedback on this story