S²VOPD: self-supervised visual on-policy distillation lifts Qwen3.5-4B from 70.7% to 77.4% on six fine-grained perception benchmarks
UC San Diego
Creates teacher-student asymmetry by subtracting information from the student through visual augmentations rather than adding privileged information to the teacher. Distilling the teacher's distribution on-policy into the student's distribution on a strongly augmented view lifts Qwen3.5-4B from 70.7% to 77.4% across six fine-grained perception benchmarks, recovering 96% of the gain from privileged-information methods without ground-truth, rewards, or a stronger teacher.
Why it matters
159 upvotes on HuggingFace Daily Papers. Pure self-supervised recipe that recovers nearly all the gain of privileged-information distillation for VLMs — useful for anyone fine-tuning open VLMs without reward models.
Importance: 4/5
HF Daily 159 upvotes
Sources
official
S²VOPD (HF Daily Papers)
media
arXiv listing