A new benchmark tests whether AI assistants can reason over years of personal photos and changing preferences rather than merely retrieve an isolated remembered fact. ReaLMem is built from authentic, multi-year visual archives paired with first-person subjective annotations.
The evaluation divides personal memory into three levels. Factual recall asks what happened; persona inference asks what the history implies about the person; predictive personalization asks the system to anticipate a useful choice or response. That progression targets a limitation of existing long-term-memory tests, which are often synthetic, text-only and focused on straightforward retrieval rather than causally connected life events.
The researchers also introduce ChronoProfiler, a module that estimates how stable a user attribute is over time and gives more weight to evidence that remains relevant. This is meant to resolve conflicts when an old preference no longer matches recent behavior. Such temporal reasoning could prevent an assistant from repeatedly applying obsolete assumptions, but real personal archives also make consent, privacy and data security central concerns. Benchmark performance does not establish that a product can safely hold someone’s life history or infer sensitive traits appropriately. ReaLMem is presented in a new arXiv preprint, so its dataset construction, representativeness and evaluation criteria still need external examination.