IBM Research benchmarks how much agentic memory each model tier actually needs

IBM Research

Research official 1 src. ~1 min

IBM Research published the second ALTK-Evolve post on calibrating agentic-memory dose. Across 8 frontier models on AppWorld (585 tasks, 9 apps), DeepSeek-V3.2 gained +9.5pp TGC with the full guideline set, gpt-oss-120b improved most via compact curated retrieval, and GLM-5 saturated — the same memory policy does not fit all capability tiers.

Why it matters

Provides an empirical model-tier calibration for agentic memory in production: full guidelines for frontier models, retrieval-curated for smaller ones, and skip for saturated ones. Practical advice for prompt caching and cost control, with code at github.com/IBM/ALTK-Evolve.

Importance: 2/5

default

Sources