2021/12/01 by S. Indrapriyadarsini, Shahrzad Mahboubi, Indrapriyadarsini, S. +7
Computer Science · Physics and Astronomy · #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Model Reduction and Neural Networks #Neural Networks and Applications #Optimization and Control (math.OC) #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.2112.01327
openalex publication_date 2021/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The Nesterov's accelerated quasi-Newton (L)NAQ method has shown to accelerate the conventional (L)BFGS quasi-Newton method using the Nesterov's accelerated gradient in several neural network (NN) applications. However, the calculation of two gradients per iteration increases the computational cost. The Momentum accelerated Quasi-Newton (MoQ) method showed that the Nesterov's accelerated gradient can be approximated as a linear combination of past gradients. This abstract extends the MoQ approximation to limited memory NAQ and evaluates the performance on a function approximation problem.