#long-horizon
- Alaya-EVOKE: From Linear-Scaling Supervision to Endless World research
- StateM: harness scaling pushes GPT-5.6 Sol to 95.3% on Terminal-Bench 2.1 and the same harness adapts to DeepSeek-V4 Flash for $15 total spend research
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work Apodex research
- Alaya-EVOKE: From Linear-Scaling Supervision to Endless World Zhejiang University research
- EchoWM: Open and Enterable Omnimodal World Models JD.com research
- AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Microsoft research
- LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks research
- Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements National University of Singapore (Zhi Zheng, Wee Sun Lee et al.) research