- Sources: primary, discussion
- Summary: The argonautlabsai fork of deltafin reports running the full 2.8 trillion parameter Kimi K3 weights on one M5 Max MacBook Pro with 128 GB by streaming experts from four external SSDs, on the engine built by the upstream gavamedia project. The fork measures about 1.00 token per second steady decode over 512 tokens and states that a 512-token prompt takes about 6.3 minutes to its first token. Its drive ladder shows one drive at about 52 percent of four-drive speed, two at about 73 percent and three at about 90 percent.
- Why it matters: The sublinear ladder shows the slowest of each layer's 16 reads sets the pace, so latency spread across drives matters more than aggregate bandwidth.
send feedback on this story