Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

National University of Singapore (Zhi Zheng, Wee Sun Lee et al.)

Research official 2 src. ~1 min

Replaces RL with evolution strategies for fine-tuning long-horizon LLM agents, needing only inference-level GPU memory and supporting trajectory-level credit assignment. On WebArena-Lite it improves a Qwen-3.5-27B agent by 6.69 points over a No-Skill baseline and wins 28 of 36 test-time automatic-heuristic-design settings.

Why it matters

96 HF upvotes (Aug 19) — offers a viable path for academic labs to fine-tune frontier-scale agentic policies on commodity GPUs, sidestepping the cost of agentic RL infrastructure.

Importance: 2/5

default

Sources