- Sources: gavamedia/deltafin, HN 49090233
- Summary: The deltafin project reports running Kimi K3 on a single M1 Max at 14.6 seconds per token. Kimi K3 is a 2.8T-parameter mixture-of-experts model with 104B active parameters, documented in the technical report this digest carried on 2026-07-28. The project names the bottleneck as disk bandwidth rather than arithmetic: a 53 GB resident spine is re-read on every token. The 14.6 seconds per token figure is for the
--full install, which needs roughly 1.7 TB of local disk on a 64 GB M1 Max. The project's --stream mode needs about 215 GB and runs at roughly 3 minutes per token. The measurements are the project's own on one machine. - Why it matters: It puts a measured cost on running a 2.8T-parameter model on one workstation, in disk footprint as well as latency, and identifies the constraint a team would have to remove to make that setup usable.
send feedback on this story