vix.ing · top · new · best · stats · spec

Adaptive First- and Second-Order Algorithms for Large-Scale Machine Learning

2021/11/29 by Sanae Lotfi, Lotfi, Sanae, Tiphaine Bonniot de Ruisselet +5
Computer Science · Decision Sciences · Engineering · Mathematics · #68T07 #90C15 #90C30 #90C53 #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #FOS: Mathematics #G.1.6 #G.3 #G.4 #I.2.6 #Machine Learning (cs.LG) #Numerical Analysis (math.NA) #Optimization and Control (math.OC) #Sparse and Compressive Sensing Techniques #Stochastic Gradient Optimization Techniques #acm:68T07 #acm:90C15 #acm:90C30 #acm:90C53 #cs.LG #cs.NA #math.NA #math.OC #msc:68T07 #msc:90C15 #msc:90C30 #msc:90C53

paper · pdf · doi:10.48550/arxiv.2111.14761

29 pages, 8 figures. arXiv admin note: text overlap with arXiv:2012.05783

arxiv created 2021/11/29 · openalex publication_date 2021/11/29 · arxiv updated 2021/11/30 · openalex created_date 2022/11/07 · openalex updated_date 2026/07/28

Abstract

In this paper, we consider both first- and second-order techniques to address continuous optimization problems arising in machine learning. In the first-order case, we propose a framework of transition from deterministic or semi-deterministic to stochastic quadratic regularization methods. We leverage the two-phase nature of stochastic optimization to propose a novel first-order algorithm with adaptive sampling and adaptive step size. In the second-order case, we propose a novel stochastic damped L-BFGS method that improves on previous algorithms in the highly nonconvex context of deep learning. Both algorithms are evaluated on well-known deep learning datasets and exhibit promising performance.

Citations

Related