PaperGym converts research papers into AI training environments by extracting evaluation criteria from a paper’s methodology and experiments rather than reusing its original research question. A two-stage approach uses rubrics as privileged context for self-teaching, followed by reinforcement learning, tested across multiple Qwen model sizes with gains over supervised fine-tuning alone. The authors release a 20,000-instance dataset, evaluation benchmarks, and trained models.
