- Sources: primary, HN discussion
- Summary: Conviva reads 3 to 5 GB Arrow IPC files from local NVMe and found mmap thrashing under concurrency: p95 rose from about 30s to over 150s, adding pods made it worse because the host page cache is shared, and a controlled test had one pod beat four pods by 41 percent at maximum load. Off-CPU analysis put the cost in futex waits at 30.9 percent and preemption at 29.3 percent against 6.9 percent in actual disk I/O, with mmap delivering 3.44 GB/s against a 21.7 GB/s fio ceiling, and moving to io_uring with O_DIRECT cut major faults from 128,957 to 3,647, a 35x reduction, while minor faults rose 8x to 8.6 million and total query time went from 13.6s to 21.8s. The post, dated 2026-09-01, attributes the regression to the design around the I/O rather than the I/O primitive, naming a single materialization thread doing submission, completion, Arrow decode and cache management.
- Why it matters: A measured negative result against the usual direction of travel, showing that swapping the I/O primitive without restructuring the consumer moves the bottleneck rather than removing it.
send feedback on this story