- Sources: primary, discussion
- Summary: The asm-hall-of-shame repository from Christopher Domas searches for the floor of single-instruction performance, with x86 entries running from nop at 1 cycle to fxrstor64 at 198,002,498,236 cycles, about 62 seconds. The top entry uses fxrstor64 to load 512 bytes of FPU, MMX and XMM state from a high-latency MMIO region in the PCIe fabric on a Ryzen 7 5800H while a fleet of hammer cores saturates the PCIe root complex with tight 4-byte reads against a different high-latency register, so the timed load queues behind that traffic, and the entries below it name microcode assists on denormal operands, split locks that assert the external bus lock, rdrand depleting the entropy pool, and wbinvd writing the full cache hierarchy back to DRAM. Stated rules require a single non-interruptible instruction, factory stock hardware, no timing of trap handlers, and normalization to base clock, with scores self-reported by the author and the ARM and RISC-V leaderboards empty.
- Why it matters: Instruction tables bound the fast path, and a measured worst case of about 62 seconds for one instruction bounds the tail that scheduling and watchdog assumptions have to survive.
send feedback on this story