SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
OpenMOSS
Repository-level benchmark of 119 tasks across 98 GitHub repos in 20 scientific domains, spanning issue-driven, expert-exploratory and engineering-integration paradigms. Best agent (Claude Code on Opus) scores under 50% pass@1; the paper catalogues four recurring failure mechanisms.
Why it matters
HuggingFace Daily Papers: 52 upvotes. Extends the SWE-bench lineage beyond software engineering into scientific computing — useful signal that top coding agents remain fragile outside mainstream stacks.
Importance: 2/5
HF Daily Paper with 52 upvotes
Sources
official
arXiv abstract page
official
HF Daily Papers entry