• Sources: primary, discussion
  • Summary: The widely cited claim is that dynamically typed languages cost fewer tokens because they omit type declarations, with a stated 2.6x gap between C and Clojure and J at 70 tokens average, which Luu notes comes from Rosetta Code tasks solvable in tens of tokens. He runs two larger evals instead, a zstd decoder implemented from the RFC without tests and a modified Pandoc ProgramBench scored against holdout tests, and the dynamic cluster lands ahead at medium effort while the ultra effort result is mixed with several static languages among the best and obscure dense languages doing poorly. He documents a defect in one cited benchmark where a test executed a ../../minigit path that did not exist, which is what caused the Rust and Haskell failures the benchmark's author read as evidence that difficult languages burden the model. A Go agent later symlinked that path to its own executable, so scoring on that broken test after the symlink measured the Go binary, and Luu notes the Rust scoring ran before it.
  • Why it matters: Language choice for agent work is being argued from token counts measured on problems too small to separate the languages, and Luu reports the only durable signal as a weak to moderate positive correlation between language popularity and both correctness and cost.

send feedback on this story