- Sources: primary, discussion
- Summary: The post, dated 2026-08-03 and surfaced on Hacker News four days later, walks a miniature Volcano-model executor summing 500 million float8 values from 1.3 s to 480 ms with 1024-row batching, to 358 ms with operator fusion, and to 135 ms with aarch64 SIMD intrinsics, against about 20 s for the same query in PostgreSQL 18.4. The disclosed setup is an AWS c8g.4xlarge Graviton4 with max_parallel_workers_per_gather set to 0, data warm in shared buffers, and the median of 5 runs. That 9.6x is measured in a standalone executor the author wrote for the post, not inside pgrust, and the claim that pgrust's query engine drove about 10x of the 300x is asserted rather than measured here. The release-level figures are separate and weaker: 10x over pgrust 0.1, 30 percent over Postgres on OLTP, and 300x over Postgres on ClickBench are asserted from the 0.2 release, are not reproduced in this post, and have no independent replication.
- Why it matters: The post states the measured cost of each step rather than a single total, so the batching and SIMD reasoning transfers to any Volcano-model executor. The post disclaims the fusion step as hardcoding an optimization for a query known in advance, which generalises only through JIT.
- Follow-up: Record independent replication of the release-level ClickBench figure, and the deferred post on JIT compilation as the general replacement for hand-written operator fusion.
send feedback on this story