- Sources: primary, discussion
- Summary: The post, dated 2026-08-12 with four Redwood Research authors and two from Anthropic, describes three benchmarks and a weighted aggregate. LMCA carries 560 position texts and 1,461 arguments with 2,140 human ratings and measures how closely a model's ratings match the researchers', ACCoRD measures logical consistency across separately asked probability and preference questions using 567 of about 14,000 generated constraints, and DTBench capabilities is 407 handcrafted decision-theory questions, weighted 60, 20, and 20 percent. Scores are as of 2026-08-10, the authors put the practical ceiling near 91, the top score is Opus 5 at 73.6 with a 95 percent confidence interval of plus or minus 2.1, and the methodology section states that Fable 5's score used Opus 5 as a fallback where Fable 5 refused, which is a caveat on that one figure.
- Why it matters: The benchmarks target the class of task where no verifiable answer exists and training feedback loops cannot be built, which is the gap the authors argue standard evaluation misses.
send feedback on this story