• Sources: preprint
  • Summary: The preprint fixes the operation sequence ahead of execution and replays it through a CUDA graph, which removes per-kernel dispatch cost from the inner loop. The reported speedups are the authors' own and are not independently reproduced.
  • Why it matters: Fixing the operation sequence ahead of time and replaying it as one graph launch removes the per-kernel dispatch overhead that had kept this workload faster on CPUs.

send feedback on this story