OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
NLP Group of The University of Hong Kong
Proposes a benchmark for judging computer-using agent trajectories with vision-language models, finds even state-of-the-art models show a systematic leniency bias that mislabels failures as successes, and releases the OS-Shepherd-100K dataset plus 9B/35B open reward models matching commercial alternatives at lower cost.
Why it matters
Second most-upvoted paper on HuggingFace Daily Papers for 2026-08-09 with 67 upvotes; exposes a concrete evaluation flaw affecting how computer-use agents are graded across the field.
Importance: 3/5
Notable HuggingFace Daily Paper (research) — Second most-upvoted paper on HuggingFace Daily Papers for 2026-08-09 with 67 upvotes; exposes a concrete evaluation flaw affecting how computer-use agents are graded across the field.