2021/06/27 by Chandak, Siddharth, Borkar, Vivek S., Dodhia, Parth · 4 citations
#FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Systems and Control (eess.SY) #electronic engineering #information engineering
paper · doi:10.48550/arxiv.2106.14308
Using a martingale concentration inequality, concentration bounds `from time n0 on' are derived for stochastic approximation algorithms with contractive maps and both martingale difference and Markov noises. These are applied to reinforcement learning algorithms, in particular to asynchronous Q-learning and TD(0).