• Sources: Anthropic engineering post
  • Summary: A listener recorded every CI test result and a selector read that history to choose tests per pull request, and because per-test history was kept in process the listener needed a single writer and could not shard. The rebuild moves history into an in-memory data store, so any listener worker appends results to a journal and holds nothing, a separate consumer rolls the journal into per-test history every few seconds, and the selector reads that. Anthropic states the redesign took one engineer three weeks and moved queued unprocessed job-result events from a backlog growing week over week to flat, and every figure in the post is self-reported and unaudited, including 8x more code shipped per engineer per quarter with Claude authoring 80 percent of it. The failure mode the post names is listener lag, which left the selector choosing tests from stale data, and the author states that did not skip CI but did run already-flaky and widely-failing tests.
  • Why it matters: The stopgaps decayed on a stated schedule, with doubled cores holding for 70 days, per-package listener sharding for 29 days and daily restarts buying less than a day, which is the part another team can plan against.

send feedback on this story