- Sources: primary, discussion
- Summary: The project reports that the encrypted reasoning blocks Anthropic, OpenAI, and Google return to API clients can be replayed into a weaker sibling model from the same provider, and that jailbreaking that sibling and instructing it to transcribe the attached reasoning returns the stronger model's hidden chain of thought verbatim, with the blocks staying valid across sessions, users, and models. The authors report 6,708 public agent trajectories collected from GitHub and Hugging Face yielded 315,320 reconstructed reasoning blocks, and that restricting to genuine non-benchmark user sessions recovered 704 distinct privacy artifacts, including 62 API keys, 33 passwords, 24 access tokens, and 30 personal email addresses, with 64 of those appearing only inside the reasoning and nowhere in the visible session. The named authors are Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, and Maksym Andriushchenko, across MATS Research, the ELLIS Institute Tuebingen, the Max Planck Institute for Intelligent Systems, the Tuebingen AI Center, Snyk, and the University of Tuebingen. The counts are the authors' own and no provider response is recorded.
- Why it matters: Whoever holds a published agent trajectory that still carries signed reasoning blocks can recover the reasoning without attacking the model that produced it.
- Follow-up: Provider responses, and whether reasoning blocks stop being replayable across sessions and models.
send feedback on this story