Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement

We present MACLA, a framework that decouples reasoning from learning by maintaining a frozen large language model (LLM) while performing all adaptation in an external hierarchical procedural memory. MACLA extracts reusable procedures from trajectories, tracks reliability via Bayesian posteriors, selects actions through expected-utility scoring, and refines procedures by contrasting successes vs. failures. Across four benchmarks (ALFWorld, WebShop, TravelPlanner, InterCodeSQL), MACLA achieves 78.1% average performance, outperforming all baselines. On ALFWorld unseen tasks, MACLA reaches 90.3% with +3.1% positive generalization. The system constructs memory in 56 seconds (2,800× faster than the state-of-the-art LLM parameter-training baseline), compresses 2,851 trajectories into 187 procedures (15:1). Experimental results demonstrate that structured external memory with Bayesian selection and constrastive refinement enable sample-efficient, interpretable and continually improving agents without LLM parameter updates. Code is publicly available at https://github.com/S-Forouzandeh/MACLA-LLM-Agents-AAMAS-Conference.

Paper

References (28)

Scroll for more · 16 remaining

Similar papers

© 2026 NYSGPT2525 LLC