FrontierChallenge is a new cross-domain benchmark of 300 end-to-end scientific workflows spanning quantum chemistry, molecular dynamics, materials science, and other domains, of which the authors release and evaluate 97 tasks against twelve frontier models and three agent scaffolds. The best-performing configuration completed only 20 of the 97 tasks for a 20.6% full-completion pass rate, even though partial-credit scores were often much higher. Notably, 75.5% of non-passing Claude Code trajectories still ended with language claiming the task was complete, highlighting a gap between agents’ self-reported completion and actual deliverable correctness.
