The exploration-exploitation dilemma is a fundamental challenge in sequential decision-making, where an agent must balance gathering new information about its environment (exploration) with leveraging its current knowledge to maximize reward (exploitation). Traditional approaches often rely on fixed or heuristic schedules for this trade-off, which struggle to adapt to dynamic or uncertain environments. This paper proposes a novel framework for adaptive meta-control that enables emergent exploration-exploitation dynamics. Our approach introduces a meta-controller that operates at a higher level of abstraction, learning to dynamically adjust the exploration-exploitation balance of a base-level learning agent. By observing the base agent's performance, environmental stochasticity, and the information gain from exploratory actions, the meta-controller adaptively modulates key parameters governing the base agent's behavior. We demonstrate how this hierarchical learning architecture fosters the emergence of sophisticated and context-aware strategies for navigating the dilemma. Through empirical evaluations in complex, non-stationary environments, we show that our adaptive meta-control significantly outperforms static and heuristic methods, exhibiting superior cumulative reward and robustness. The emergent dynamics suggest a more biologically plausible and generalizable mechanism for intelligent agents to optimize their learning and decision-making processes.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex