Dynamic Important Example Mining (DIEM) adaptively selects and weights training data throughout reinforcement fine-tuning, combining a gradient-alignment importance estimator that identifies each sample’s contribution to policy improvement with a constrained batch-reweighting scheme that optimizes aggregate utility while keeping optimization stable. The authors report this dynamic approach outperforms static and other adaptive baselines across reasoning benchmarks.