- Sources: primary, discussion
- Summary: The write-up publishes the rig rather than only a score: supervised fine-tuning distilled from about five hundred frontier-model trajectories, then a GRPO variant whose reward is measured execution time against Postgres's own plan, with vLLM and the trainer on a rented two-GPU node and four Postgres containers isolated to limit page-cache contention. The starting 4B model could not produce a plan at all for 99 of the 113 queries. The reported result is a 1.81 times geometric-mean speedup and a 44.7 percent summed latency reduction on the Join Order and Cardinality Estimation benchmarks, measured by the author alone.
- Why it matters: The reward is wall-clock execution time against the planner's own choice, which is a reproducible target rather than a proxy metric.
send feedback on this story