GPU-CFR compiles counterfactual regret minimization to static dataflow and CUDA graph replay
- Sources: preprint
- Summary: The preprint fixes the operation sequence ahead of execution and replays it through a CUDA graph, which removes per-kernel dispatch cost from the inner loop. The reported speedups are the authors' own and are not independently reproduced.
- Why it matters: Fixing the operation sequence ahead of time and replaying it as one graph launch removes the per-kernel dispatch overhead that had kept this workload faster on CPUs.