- Sources: watch
- Channel: AI Engineer (2026-09-09, 3,667 views, 5.0 over 20 ratings)
- Summary: Laurie Voss replicated the year-old IFScale result on the models from the original paper, then pointed the same test at current frontier models, which scored 100 percent immediately. He raised the benchmark from 500 words to 10,000, and the boundary now sits near 2,000 instructions and closer to 5,000 for the best model. Failure modes differ per model, from silent forgetting to safety-classifier refusal to politely declining after writing five thousand words.
- Why it matters: Fitting rules into a small budget is no longer the binding constraint, and verifying that the model obeyed replaces compressing the rules.
send feedback on this story