SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

OpenMOSS

Research official + media 2 src. ~1 min

Repository-level benchmark of 119 tasks across 98 GitHub repos in 20 scientific domains, spanning issue-driven, expert-exploratory and engineering-integration paradigms. Best agent (Claude Code on Opus) scores under 50% pass@1; the paper catalogues four recurring failure mechanisms.

Why it matters

HuggingFace Daily Papers: 52 upvotes. Extends the SWE-bench lineage beyond software engineering into scientific computing — useful signal that top coding agents remain fragile outside mainstream stacks.

Importance: 2/5

HF Daily Paper with 52 upvotes

Sources