- Sources: PyTorch blog, HN 49048689
- Summary: A joint AMD and Meta post dated 2026-07-06, surfaced on Hacker News on 2026-07-25, describes porting PyTorch Monarch to ROCm. Monarch orchestrates a GPU cluster from a single Python program using an actor runtime, a process mesh abstraction, and supervision-tree fault handling. The port covers three paths: collective communications converted from CUDA to HIP with
hipify_torch and linked against RCCL, GPU memory management routed through HIP driver calls, and the libibverbs RDMA path kept while GPU-side bindings move to HIP. A Rust compatibility shim re-exports HIP symbols under CUDA names. The authors report all 1,171 tests passing on ROCm 7.0 and above, a 16-node MI300 SLURM run of 128 GPUs training Llama 3 8B with RCCL failures injected every 180 seconds and no full restart, and a 32-node MI355 Kubernetes run of 256 GPUs where participant count fluctuated between 30 and 32 during recovery while loss fell from 12 to about 4. They list extended NIC support, rejoin reload latency, and overlapping recovery with compute as open work. - Why it matters: Fault-tolerant single-controller training on a non-CUDA stack at 256 GPUs is a concrete data point on how far the ROCm path has come for large training jobs.
- Follow-up: Watch for independent reproduction of the fault-injection runs and for whether other pre-training and reinforcement-learning frameworks land on ROCm.
send feedback on this story