ByteDance Seed researchers present DiffusionOPSD, an on-policy self-distillation framework that converts image-level rewards into intermediate training targets for diffusion models. A frozen behavior policy generates trajectories while reward gradients construct bounded targets around anchor predictions, and the trainable policy fits these targets before an exponential moving average refreshes the behavior policy. Testing on SD 3.5-M and Z-Image-Turbo showed superior performance in 19 of 20 settings, with up to 44% improvement over competing methods and 40-63% lower training GPU hours.
