- Sources: primary, discussion
- Summary: The post describes compiling Rust's
core::simd to GPU warp instructions, mapping SIMD lanes onto warp lanes, reductions onto warp shuffles, and mask queries onto vote and ballot instructions. It states the mapping is one to one only when the vector width matches the hardware warp width, which is 32 lanes on NVIDIA. - Why it matters: Portable SIMD code can target a GPU without a separate kernel language, at a cost that depends on matching the warp width.
send feedback on this story