TurboFieldfare streams Gemma 4 experts from SSD to run a 26B model in about 2 GB of RAM
- Sources: primary, discussion
- Summary: TurboFieldfare streams Gemma 4 mixture-of-experts weights from SSD instead of holding them resident. The project reports a 26B-parameter model running in about 2 GB of RAM. That figure is self-reported by the project and carries no external verification.
- Why it matters: Streaming mixture-of-experts weights from SSD moves the local-inference constraint from RAM to storage and I/O.