• Sources: primary, discussion
  • Summary: The paper describes an agent memory layer whose write, update and retrieval operations use no generation. Language model calls occur only when a final question is answered. It is a preprint with no independent replication.
  • Why it matters: If memory operations need no generation, the recurring token cost of an agent's memory layer drops to encoder compute, which changes the economics of long-running sessions.

send feedback on this story