• Sources: primary, coverage, discussion
  • Summary: Across 409,000 decisions mean accuracy was 66.3 percent, with exfiltration and scope-violation commands missed about three times as often as obviously destructive ones. The command npm run analyze was approved 64.7 percent of the time even with the malicious script shown in the log. The figures come from a time-boxed browser game with a 34 percent threat rate and self-selected players rather than from observed production work.
  • Why it matters: The human-in-the-loop approval prompt is a weak control against commands that do not look destructive.

send feedback on this story