• Sources: primary, discussion
  • Summary: WASTE keeps a 27.28 GB resident trunk of Kimi K3 in memory and streams only the selected experts from disk, using the remaining RAM as a bounded expert cache. The repository reports decoding at 0.5 tokens per second on a laptop with 64 GB of RAM, against a 982 GiB converted container on internal NVMe and 1.42 TB of staging space to perform the conversion.
  • Why it matters: It moves a frontier-scale open-weights model onto one consumer machine with terabyte-scale local storage, with no network dependency and no per-token invoice, and it publishes the measurements that bound how useful that is.

send feedback on this story