Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
UC San Diego (Yunhao Yang, Nuno Vasconcelos, Yijiang Li et al.)
Multiple decoupled models share no parameters but are jointly optimized via RL with peer-derived rewards; cohort diversity across model families, sizes, and rephrased samples cuts correlated-error feedback loops. Yields 3.0–8.6% gains on seven LLM benchmarks and 2.3–7.2% on four VLM benchmarks without any ground-truth labels.
Why it matters
40 HF upvotes (Aug 20) — a label-free recipe that matches or beats supervised reasoning training on text and vision-language tasks, which would lower the cost barrier for reasoning-model post-training.
Importance: 2/5
default
Sources
official
arXiv — Co-RL paper
official
HuggingFace Daily Papers — Co-RL