Researchers from the University of Chinese Academy of Sciences and Meituan introduced an evaluation framework that assesses AI research agents beyond final task scores, testing seven frontier language models across 36 long-horizon tasks from AutoLab. The study measured agents on solution framing, execution, and feedback control, finding that current systems function largely as engineering optimizers rather than autonomous researchers, with genuine methodological novelty appearing in only 1.2% of solutions. The researchers also found that accumulated experience can both help and hurt subsequent performance, while agent harness design mainly affects consistency rather than peak capability.
