IBM Research’s ALTK-Evolve framework examines how AI agents benefit from memory depending on their capability level, testing the approach across eight models. The study found that strong models with spare capacity gain most from full guideline sets, weaker models perform better with selective retrieval of task-relevant guidelines, and already-saturated models show no measurable improvement from added memory. Curated retrieval improved performance while minimizing token overhead, with gpt-oss-120b achieving a 16.1 percentage point gain in task completion using only 5% more tokens, suggesting optimal memory strategies should prioritize calibration over simple accumulation.