Researchers introduce MemTrapBench, a benchmark showing that retrieved memories can actually harm large language model task performance even when the retrieved information is accurate and relevant. The benchmark measures two failure modes, termed Reasoning Fixation and Belief Distortion, across multiple model families and memory frameworks. The study found that all evaluated memory strategies underperformed a no-memory baseline, with even the strongest methods suffering performance drops of more than 10%, and proposes an inference-time technique called AdaptiveMem to mitigate the effect.
