On-Policy Self-Distillation without Any Supervision

Research official 2 src. ~1 min

U-OPSD trains a language model on consensus pseudo-solutions built from its own majority-voted generations, correcting confident mistakes without ground-truth labels, teacher models, or environment feedback.

Why it matters

187 upvotes on HuggingFace Daily Papers; reports 8.5-10.7% gains over base models on math benchmarks (AIME24/25, HMMT25, MATH500, AMC23), competitive with supervised self-improvement methods.

Importance: 3/5

HuggingFace Daily Papers with 187 upvotes (>=100 bump threshold applied).

Sources