Princeton University researchers present Recuris, an architecture for long-horizon agent tasks built on two coupled memory systems: a Working Memory that tracks task progress and guides skill selection from an Experiential Memory, and a Meta-Agent that iteratively refines Skill Memory based on execution evidence. Evaluated across four benchmarks and ten models, the approach delivered gains including +17.8 points to GPT-5.6 Sol on tau-bench and raised Claude Opus 5 to 87.9% task success, with improvements expanding to +32.2 points on extended-horizon tasks.
