A new benchmark called MobileMem evaluates AI agents’ ability to provide persistent personal assistance by synthesizing a full year of simulated mobile-device usage into coherent multimodal trajectories. The benchmark tests agents on multi-hop and temporal reasoning, knowledge updating, and implicit preference inference grounded in realistic phone interactions rather than isolated facts, aiming to shift memory-system evaluation toward what its authors call experiential intelligence for continuous personal learning.
