2024/05/27 by Junwen Qiu, Qiu, Junwen, Bohao Ma +3 · 1 citation
Computer Science · Engineering · Physics and Astronomy · #Advanced Thermodynamics and Statistical Mechanics #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Matrix Theory and Algorithms #Optimization and Control (math.OC) #Sparse and Compressive Sensing Techniques
paper · pdf · doi:10.48550/arxiv.2405.16954
openalex publication_date 2024/05/27 · openalex created_date 2024/05/29 · openalex updated_date 2026/07/28
The stochastic gradient descent method with momentum (SGDM) is a common approach for solving large-scale and stochastic optimization problems. Despite its popularity, the convergence behavior of SGDM remains less understood in nonconvex scenarios. This is primarily due to the absence of a sufficient descent property and challenges in simultaneously controlling the momentum and stochastic errors in an almost sure sense. To address these challenges, we investigate the behavior of SGDM over specific time windows, rather than examining the descent of consecutive iterates as in traditional studies. This time window-based approach simplifies the convergence analysis and enables us to establish the iterate convergence result for SGDM under the Łojasiewicz property. We further provide local convergence rates which depend on the underlying Łojasiewicz exponent and the utilized step size schemes.