Equipping an AI agent with agentic memory—the ability to distill lessons from past work and re-inject them into context—is often viewed as a simple equation: more experience equals better performance. However, new research using the ALTK-Evolve framework suggests that memory is not a feature to be switched on, but a "dosage" that requires precise calibration based on the model's tier.
Learning Around the Model, Not Inside It
ALTK-Evolve allows an agent to learn from its own past trajectories by distilling reusable guidelines and injecting them back at inference time. This process involves no weight updates and no human annotation. The loop is straightforward: the agent attempts tasks, the system extracts behavioral guidelines from both successes and failures, and then consolidates them into a reusable set provided to the model.
The Three Patterns of Memory Dosage
The research, conducted across eight models using the AppWorld benchmark, revealed three distinct patterns:
- Strong models with headroom: Systems like DeepSeek-V3.2 (671B MoE) thrive on the full guideline set, seeing a +9.5 percentage point jump in task completion.
- Smaller or weaker models: These models can be overwhelmed by large guideline sets. For instance, gpt-oss-120b (117B MoE) achieved its best gains (+16.1pp) using "curated retrieval"—a compact core of task-relevant guidelines.
- Saturated models: Models like GLM-5 (745B MoE) showed no measurable gain, suggesting they may already be near their performance ceiling for the tested tasks.
Efficiency and Cost-Effectiveness
The study found that the cheapest memory strategy can also be the best. Curated retrieval kept costs near the baseline, with gpt-oss-120b gaining significant accuracy at only a +5% token cost. In production, prompt caching is a key efficiency lever, as it allows the static portion of the guideline set to remain stable across steps, cutting effective costs substantially. Even models near the ceiling on task completion, such as GPT-5.5 and Opus, showed gains of +7.2 and +7.1pp respectively in the stricter Scenario Goal Completion (SGC) metric when provided with memory.