EnvHarness: Awakening Static Worlds for Agent Learning
An 'ActionableEnv' wrapper plus 'EnvRigger' that programmatically reshape static LLM-agent benchmarks via Setup/Rule/Link plugins, targeting specific agent weaknesses without rebuilding verifiers; across five benchmarks in four domains, lifts SWE-bench Verified resolution from 52.13% to 54.79% and trims average steps per episode by 9.8%.
Почему это важно
HF Daily Paper with 248 upvotes; addresses the benchmark-saturation problem for LLM-agent RL
Важность: 3/5
HF Daily Paper with 248 upvotes
Источники
официальный
google-research/envharness
официальный
Project page