- Sources: primary, discussion
- Summary: Cursor open-sourced Mixture-of-Kittens, a production mixture-of-experts training megakernel that fuses all MoE communication and computation into a single deterministic kernel and powers Composer training across tens of thousands of GPUs. The post reports 2.37x higher MXFP8 forward throughput than the fastest public baseline and 1.41x end-to-end tokens per second. A signalling microbenchmark in the post shows dispatch latency falling from 103 to 18 microseconds when receivers pull tokens instead of senders pushing them.
- Why it matters: The result inverts the usual assumption that push saturates NVLink better, and the reason is signalling rather than bandwidth, because pull-based dispatch removes cross-GPU completion signals entirely.
send feedback on this story