Researchers from Peking University and BUPT identified a failure mode in sequential reinforcement learning with verifiable rewards, where optimizing for one objective makes successful behaviors needed for a different objective too rare to sample later. Using controlled interventions on IFEval, AIME, and MATH, they showed that math-focused RLVR training improves instruction-following on average but shrinks the share of prompts with discoverable solutions under repeated sampling, while instruction-following-focused training pushes math responses toward direct answers instead of step-by-step reasoning. The effects were traced to distributional shifts concentrated in the openings of model responses.
