Demystifying Agent Skills: Why They Work — Until They Don't

UC San Diego (Zhiyuan Jiang, Mengdi Wang, Yijiang Li et al.)

Research official 2 src. ~1 min

Controlled experiments across benchmarks, agent harnesses, and LLMs show skills beat Workflow Memory by 6.06 points but retrieval precision collapses from 29.6% to 3.3% as skill pools grow from 5 to 100. Procedural anchoring accounts for 65.7% of skill wins versus 4.5% from explicit knowledge injection.

Why it matters

Top paper on HF Daily Papers for Aug 19 with 149 upvotes — gives the first systematic ablation of when and why agent-skill libraries actually help, and surfaces the scaling cliff practitioners hit.

Importance: 3/5

HF Daily ≥100 upvotes

Sources