Understanding the Origin of Information-Seeking Exploration in Probabilistic Objectives for Control

The exploration-exploitation trade-off is central to the description of\nadaptive behaviour in fields ranging from machine learning, to biology, to\neconomics. While many approaches have been taken, one approach to solving this\ntrade-off has been to equip or propose that agents possess an intrinsic\n'exploratory drive' which is often implemented in terms of maximizing the\nagents information gain about the world -- an approach which has been widely\nstudied in machine learning and cognitive science. In this paper we\nmathematically investigate the nature and meaning of such approaches and\ndemonstrate that this combination of utility maximizing and information-seeking\nbehaviour arises from the minimization of an entirely difference class of\nobjectives we call divergence objectives. We propose a dichotomy in the\nobjective functions underlying adaptive behaviour between \\emph{evidence}\nobjectives, which correspond to well-known reward or utility maximizing\nobjectives in the literature, and \\emph{divergence} objectives which instead\nseek to minimize the divergence between the agent's expected and desired\nfutures, and argue that this new class of divergence objectives could form the\nmathematical foundation for a much richer understanding of the exploratory\ncomponents of adaptive and intelligent action, beyond simply greedy utility\nmaximization.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC