- Sources: preprint
- Summary: During decode each request contributes one new token, so dense projection GEMMs become small-row matrix multiplications. The bfloat16 GMMA path on Hopper executes these in fixed 64-row fragments, so small-batch decode fills only a fraction of each fragment with real token rows while the utilization counter still reads high. The authors profile vLLM with FlashAttention-3 and cuBLASLt on an H100 NVL and replace the single number with eight views, each pinned to an Nsight Compute counter or an explicit formula, in results that are the authors' own and are not independently reproduced.
- Why it matters: A serving team reading high SM utilization during decode can be reading fragment occupancy rather than throughput, which points capacity planning at the wrong bottleneck.
send feedback on this story