LoopArena is a benchmark that evaluates how effectively one AI model can direct a separate coding agent through extended, multi-round tasks. A ‘Controller’ model receives structured summaries after each coding round and decides what the fixed ‘Worker’ agent should do next or whether to stop. The benchmark uses three evaluation settings that isolate loop-decision quality, repeated control over task segments, and full end-to-end task completion, with the best-performing controller reaching only a 24.69 percent success rate on full tasks, indicating substantial room for improvement in long-horizon loop control.