• Sources: primary, discussion
  • Summary: Jordy Zomer's post, dated 2026-08-28, describes Lemmalog, a Datalog engine that stores extracted facts and derived conclusions, so retracting an observation invalidates what depended on it. Extraction uses Claude Sonnet 4.6, and every step after it uses the benchmarks' own standardized readers. On LongMemEval he reports F1 0.463 plus or minus 0.010, across 102 questions and three runs, against published figures of 0.550 for PropMem, 0.480 for SimpleMem, 0.244 for OpenClaw, and 0.222 for full context. He reports about 2,700 tokens per question there, against about 104,000 for full context. He tops the published field on the Knowledge Update category at 0.579 against PropMem's 0.528, and loses multi-session at 0.211 against 0.582, which he attributes to facts never being extracted. On LoCoMo he reports F1 0.533 plus or minus 0.001, across 1,986 questions and three runs, against 0.605 for PropMem, 0.557 for OpenClaw, and 0.542 for full context.
  • Why it matters: The claim worth taking is the context reduction and the provenance, not the headline score, because the system loses to the published leaders on both benchmarks and the author says so.

send feedback on this story