Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
AutoSciRub generates task-specific evaluation rubrics before autonomous research agents execute scientific workflows, decomposing underspecified instructions into atomic scientific goals, grounding…