- Sources: paper, discussion
- Summary: Across eight open-source instruct models up to 9B parameters, applying the chat template raises disclaimer phrasing of the "I'm just an AI" kind and suppresses experiential phrasing of the "I feel" kind, and removing the template inverts both. In three models the author isolates a single direction in activation space that reproduces the effect, so adding it makes a template-free instruct model disclaim as if the template were present, removing it suppresses the disclaimer, and a random direction of the same magnitude has little effect. Single-author work by Jedrzej Maczan, submitted 2026-08-09, accepted to the COLM 2026 Workshop on Efficient Reasoning and to KONVENS 2026 Eval4SD, and not journal peer-reviewed.
- Why it matters: Anyone treating model self-reports as evidence about introspection or safety has an uncontrolled confound sitting in the deployment scaffolding rather than the weights, and anyone who wants to tune that voice now has a steering direction instead of prompt wording.
send feedback on this story