- Sources: primary, discussion
- Summary: The paper reports that substituting a closed-form symbolic structure for a network's representation process leaves the network's behavior largely unchanged. The authors report the approximation holds for large language models across arithmetic, logic, computer code and language. They report that precise interventions on the identified structures change model behavior in targeted ways.
- Why it matters: An approximation that survives substitution and supports targeted intervention is a handle on steering model behavior rather than only on describing it.
- Follow-up: Track whether independent groups reproduce the substitution result on other model families.
send feedback on this story