2021/12/11 by Francesco De Lellis, De Lellis, F., Marco Coraggio +7
Computer Science · Engineering · #Advanced Control Systems Optimization #FOS: Computer and information sciences #FOS: Electrical engineering #FOS: Mathematics #Formal Methods in Verification #Machine Learning (cs.LG) #Optimization and Control (math.OC) #Reinforcement Learning in Robotics #Systems and Control (eess.SY) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2112.06018
openalex publication_date 2021/12/11 · openalex created_date 2021/12/31 · openalex updated_date 2026/07/28
We present an architecture where a feedback controller derived on an approximate model of the environment assists the learning process to enhance its data efficiency. This architecture, which we term as Control-Tutored Q-learning (CTQL), is presented in two alternative flavours. The former is based on defining the reward function so that a Boolean condition can be used to determine when the control tutor policy is adopted, while the latter, termed as probabilistic CTQL (pCTQL), is instead based on executing calls to the tutor with a certain probability during learning. Both approaches are validated, and thoroughly benchmarked against Q-Learning, by considering the stabilization of an inverted pendulum as defined in OpenAI Gym as a representative problem.