S²VOPD: self-supervised visual on-policy distillation lifts Qwen3.5-4B from 70.7% to 77.4% on six fine-grained perception benchmarks

UC San Diego

Research official + media 2 src. ~1 min

Creates teacher-student asymmetry by subtracting information from the student through visual augmentations rather than adding privileged information to the teacher. Distilling the teacher's distribution on-policy into the student's distribution on a strongly augmented view lifts Qwen3.5-4B from 70.7% to 77.4% across six fine-grained perception benchmarks, recovering 96% of the gain from privileged-information methods without ground-truth, rewards, or a stronger teacher.

Why it matters

159 upvotes on HuggingFace Daily Papers. Pure self-supervised recipe that recovers nearly all the gain of privileged-information distillation for VLMs — useful for anyone fine-tuning open VLMs without reward models.

Importance: 4/5

HF Daily 159 upvotes

Sources