• Sources: doubleword.ai post, HN discussion
  • Summary: The post traces a single global load through the memory path of an RTX 4090 and puts a measured latency on each level, 15.4 ns when it hits L1, 127.4 ns at L2 and 255.4 ns when it reaches DRAM. It also reverse engineers two mappings Nvidia does not document, the function that selects an L1 set and the hash that assigns an address to an L2 slice. The probes and the clock settings used to take the measurements are published alongside the numbers.
  • Why it matters: The post publishes its probes and clock settings, so the latencies and the slice mapping can be reproduced on another card rather than taken from vendor documentation that does not cover them.

send feedback on this story