- Sources: project page, arXiv 2608.13545, HN discussion
- Summary: A group from MPI for Intelligent Systems, the ELLIS Institute Tuebingen, and ETH Zuerich built an 88B-token corpus filtered from FineWeb-Edu against Common Core K-5 standards, excluding concepts, facts, and vocabulary taught above Grade 5, then trained models at 0.6B, 1.3B, and 5B from scratch on it against matched unfiltered controls sharing architecture, token count, and recipe. The reported result is that scaling, SFT plus GRPO post-training, and in-context learning all amplify in-scope ability while none meaningfully improves out-of-scope performance, including when the post-training uses out-of-scope data, which places the effective capability ceiling in the pretraining filter. Model checkpoints and controls are published. This is a preprint and the finding is not independently reproduced.
- Why it matters: The controlled boundary separates a capability a model acquired from one a prompt merely elicited, which normal pretraining corpora make untestable.
- Follow-up: Watch for peer review or for an independent replication using the published checkpoints.
send feedback on this story