2021/03/11 by Beren Millidge, Anil C. Seth, Millidge, Beren +4 · 1 citation
Computer Science · Economics, Econometrics and Finance · #Artificial Intelligence (cs.AI) #Complex Systems and Time Series Analysis #Computability, Logic, AI Algorithms #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.2103.06859
openalex publication_date 2021/03/11 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
The exploration-exploitation trade-off is central to the description of\nadaptive behaviour in fields ranging from machine learning, to biology, to\neconomics. While many approaches have been taken, one approach to solving this\ntrade-off has been to equip or propose that agents possess an intrinsic\n'exploratory drive' which is often implemented in terms of maximizing the\nagents information gain about the world -- an approach which has been widely\nstudied in machine learning and cognitive science. In this paper we\nmathematically investigate the nature and meaning of such approaches and\ndemonstrate that this combination of utility maximizing and information-seeking\nbehaviour arises from the minimization of an entirely difference class of\nobjectives we call divergence objectives. We propose a dichotomy in the\nobjective functions underlying adaptive behaviour between \evidence\nobjectives, which correspond to well-known reward or utility maximizing\nobjectives in the literature, and \divergence objectives which instead\nseek to minimize the divergence between the agent's expected and desired\nfutures, and argue that this new class of divergence objectives could form the\nmathematical foundation for a much richer understanding of the exploratory\ncomponents of adaptive and intelligent action, beyond simply greedy utility\nmaximization.\n