Researchers built AutoResearchEval, a benchmark of 100 frontier research tasks spanning seven scientific domains and the full research lifecycle, and evaluated eight different agent-harness and model combinations across 800 trajectories. From this they compiled the AutoResearch Failure Taxonomy (ARFT), cataloguing 45 empirically grounded failure patterns. The study concludes that current agents lack a metacognitive loop needed to verify their own outputs, revise mistakes, or question their approach, a limitation observed across all tested models regardless of underlying strength or scaffolding.