• Sources: watch
  • Channel: AI Engineer (2026-07-23, 16,536 views, 5.0 over 733 ratings)
  • Summary: Dex Horthy of HumanLayer recounts a July 2025 experiment running an agent software factory in which nobody read the generated code, and the failure that followed: a defect no amount of prompting fixed, in a codebase he had stopped reading months earlier. He argues this is a model-training problem rather than a harness or scale problem, because coding models are reinforced on whether a test passes without breaking another and nothing in that reward penalizes architecture whose cost arrives months later. He credits Claude Code's traction to being the first model trained against the harness it ships in, and proposes planning up front through product review, system architecture, program design down to types and call graphs, then vertical slices.
  • Why it matters: It names a specific reason agent output degrades over months that neither more tokens nor a different harness addresses.

send feedback on this story