Researchers proposed S²VOPD, a self-supervised on-policy distillation method that creates teacher-student asymmetry by degrading the student’s input rather than augmenting the teacher’s. The technique distills a teacher model’s predictions on original images into a student trained on heavily augmented views, without requiring ground-truth labels or a separate teacher model. Across six fine-grained visual perception benchmarks, the method improved Qwen3.5-4B’s accuracy from 70.7% to 77.4%, surpassing larger open-source models.
