- Sources: primary, discussion
- Summary: The post benchmarks AmpereOne, Cortex-X4, Cortex-X925, Oryon-3 and Apple M1, separating acquire-load and release-store on ARMv8.0-a, the LRCPC instructions mandatory since ARMv8.3, and Apple's hardware TSO mode. Unaligned x86 accesses force FEX to patch its JIT output back to a plain load-store wrapped in a data memory barrier, which costs roughly half the throughput on Cortex-X4 and Cortex-X925 and more on Oryon-3 stores, while the M1 with TSO enabled shows aligned and unaligned within about 5 percent. The stated conclusion is that a hardware TSO toggle beats successive LRCPC extension versions.
- Why it matters: The measured gap sets what x86-on-ARM emulation costs per core design, and points the fix at a silicon feature rather than at the JIT.
send feedback on this story