A new arXiv paper introduces Supersede, a benchmark and training setup for the memory-update gap in LLM agents. The work focuses on cases where an agent keeps relying on outdated information after new evidence should supersede it.

The problem is central to long-running agents. Systems that operate across sessions need to revise assumptions, not just retrieve prior context more effectively.

By turning memory updates into a measurable failure mode, the paper adds another piece to the agent reliability toolkit, especially for assistants that manage evolving tasks or records.