Researchers propose a game-theoretic framework for understanding and improving multi-agent LLM systems in which an orchestrator decomposes tasks for a team of worker agents that improve through textual reflection. The work models orchestrator-worker interaction as a bilevel coordination game and introduces Stochastic Reflective Memory Ascent (SRMA), a method that accepts a candidate memory update only after a grounded evaluation confirms it reduces risk. On 500 SWE-bench instances, the complete Kimi-based system resolves 72.2% of tasks, compared to 70.8% for a public mini-SWE-agent reference.
