2013/05/09 by Shenghuo Zhu, Zhu, Shenghuo
Computer Science · Engineering · Mathematics · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Markov Chains and Monte Carlo Methods #Sparse and Compressive Sensing Techniques #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.1305.2218
openalex publication_date 2013/05/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
With a weighting scheme proportional to t, a traditional stochastic gradient descent (SGD) algorithm achieves a high probability convergence rate of O(κ/T) for strongly convex functions, instead of O(κ ln(T)/T). We also prove that an accelerated SGD algorithm also achieves a rate of O(κ/T).