ASI-Bench: At the Dawn of Artificial Superintelligence

42-author consortium (lead: Junwei Zhou; incl. Chi Wang, Yilun Hao, Yuantao Zhai)

Research official 2 src. ~1 min

60 project-level research tasks across 11 scientific domains, built with 31,000+ human hours and progressively less human guidance. Across 18 agent-model configurations average scores drop from 50.91 (full guidance) to 29.10 (method specified) to 26.62 (agent-determined method).

Why it matters

55 HF upvotes (Aug 19) — the first benchmark to jointly measure innovative exploration and autonomous scientific execution, with a deliberately public submission portal so the suite can grow with the field.

Importance: 2/5

default

Sources