2023/12/27 by Dylan J. Foster, Alexander Rakhlin, Foster, Dylan J. +1 · 2 voices · 6 citations
Computer Science · Engineering · Mathematics · #Advanced Control Systems Optimization #Artificial intelligence #Bayesian inference #Bayesian probability #COVID-19 epidemiological studies #Computer science #Data Stream Mining Techniques #Dilemma #Engineering #Epistemology #Frequentist inference #Function (biology) #Machine learning #Parallels #Perspective (graphical) #Reinforcement learning #Theme (computing) #cs.LG #math.OC #math.ST #stat.ML
paper · pdf · doi:10.48550/arxiv.2312.16730
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/12/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
These lecture notes give a statistical perspective on the foundations of reinforcement learning and interactive decision making. We present a unifying framework for addressing the exploration-exploitation dilemma using frequentist and Bayesian approaches, with connections and parallels between supervised learning/estimation and decision making as an overarching theme. Special attention is paid to function approximation and flexible model classes such as neural networks. Topics covered include multi-armed and contextual bandits, structured bandits, and reinforcement learning with high-dimensional feedback.