AutoSciRub generates task-specific evaluation rubrics before autonomous research agents execute scientific workflows, decomposing underspecified instructions into atomic scientific goals, grounding them in literature and data, and producing verifiable criteria that guide execution and iterative revision. Tested on ResearchClawBench and AstaBench E2E Discovery, the approach yields consistent improvements across multiple language-model configurations and agent harnesses.
