#evals
- GigaChat 3.5 Ultra passes professional-retraining information security exam, scoring 14% above pass threshold Sber research
- LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks research
- Cline open-sources Terminal-Bench-based evals for open-weight coding agents Cline tools
- Demystifying Agent Skills: Why They Work — Until They Don't UC San Diego (Zhiyuan Jiang, Mengdi Wang, Yijiang Li et al.) research
- FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis University of Science and Technology of China research
- Anthropic commits $5M to independent research on AI's impact on wellbeing Anthropic research
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? OpenMOSS research
- Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains research
- ASI-Bench: At the Dawn of Artificial Superintelligence 42-author consortium (lead: Junwei Zhou; incl. Chi Wang, Yilun Hao, Yuantao Zhai) research