- Sources: Box2D blog, HN 49013464
- Summary: Erin Catto published a post on 2026-07-18 on applying wide SIMD to convex-hull collision in Box3D. He distinguishes wide SIMD, which processes several work units at once, from narrow SIMD over the components of a single 3D vector, and notes that the separating axis test is quadratic in edge count, so a 32-point hull against another 32-point hull needs 7,921 edge-edge combinations against 144 for box against box. The optimization tests one edge of the first hull against four edges of the second at a time. On a 500-step convex pile benchmark on an AMD 7950X, single-threaded times are 40,706 ms scalar, 17,337 ms with SSE2, and 15,762 ms with an AVX2-lite path. Catto notes the gains apply to complex hulls and barely move simple box-box cases.
- Why it matters: The measured split between SSE2 and AVX2 on the same workload is a useful reminder that most of the win here comes from restructuring the loop rather than from the widest available instruction set.
send feedback on this story