• Sources: write-up, HN discussion
  • Summary: Kimi K3 at 2.8T parameters needs over 1.5 TB of VRAM before a KV cache, which rules out a single B200 node and leaves the MI355X at 288 GB per GPU as the reason the comparison exists. Both blockers the post reports were framework defects rather than missing kernels: the sglang ROCm build leaves top_k_renorm_prob undefined so the speculative-decode verifier takes the scheduler down with a NameError, fixed with a sort, a masked fill, and a divide in PyTorch, and the fast AITER MLA prefill kernel would not load because K3 at TP8 gives 12 attention heads per rank against a path built for 4, 8, or multiples of 16, fixed by zero-padding to 16 and extracting 12. The post reports the prefill fix moving a 172k-token cold prefill from the Triton fallback at 4,000 to 7,000 tokens per second to about 13,000, or roughly 51 seconds on MI355X against 23 on a B300. This is a vendor post from an inference provider, the throughput figures are unreproduced, and its performance-per-dollar claim of 48 against 33 tokens per second per dollar rests on its own stated rates of 2.50, 6.00, and 4.25 US dollars per GPU-hour.
  • Why it matters: Both blockers to serving a 2.8T-parameter model on ROCm were framework defects with small source-level fixes, so the gap was tooling rather than hardware capability.
  • Follow-up: Track whether an independent run reproduces the 952 tokens per second per node figure on eight MI355X.

send feedback on this story