• Sources: Dylan Castillo post, HN 49010129
  • Summary: A widely discussed 2026-07-22 essay argues that labs increasingly optimize for informal public evals, using Simon Willison's "draw a pelican on a bicycle" SVG test as the example, so improvements on such tests may reflect targeted training rather than general capability. The piece is opinion and offers no controlled measurement.
  • Why it matters: It restates the benchmark-contamination problem for the ad hoc evals practitioners use to compare models, arguing they degrade once a lab optimizes for them.

send feedback on this story