• Sources: primary, discussion
  • Summary: Fernando Simoes measured Turso's two I/O backends on TPC-H Q6 over a 1.2 GiB database in a post dated 2026-08-31, where the io_uring backend opens the database with O_DIRECT and no buffered option, removing kernel readahead. With prefetch off the post counts 195,207 submitted SQEs producing roughly 196,000 device requests at an average read request size of 4.37 KiB with only 140 of 195,516 bios merged, while a 32-page readahead window gives 218,212 SQEs, roughly 16,300 device requests at 56.53 KiB average size, and 202,539 of 218,493 bios merged, the stated mechanism being that the block layer can only join adjacent requests queued at the same time so a single in-flight SQE has nothing to merge against. The post separately measures sqpoll, where system time of 8.46s exceeds wall time of 8.22s with the kernel io_sq_thread taking 65 percent of cycles against 8.62s wall and 1.27s system without SQ polling, and records the syscall pread backend finishing Q6 in 3.02s against 8.55s for io_uring with 5.889 M cache misses at 7.691 percent against 10.666 M at 13.255 percent, with the author labelling the O_DIRECT page-cache explanation for that gap as a hypothesis he did not trace.
  • Why it matters: Opening with O_DIRECT moves request merging into the application, and unmerged 4 KiB requests are what leave an io_uring path slower than plain pread on the same box.

send feedback on this story