2019/09/08 by S. Indrapriyadarsini, Indrapriyadarsini, S., Shahrzad Mahboubi +5
Computer Science · Physics and Astronomy · #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and ELM #Model Reduction and Neural Networks #Neural Networks and Applications
paper · pdf · doi:10.48550/arxiv.1909.03620
openalex publication_date 2019/09/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
A common problem in training neural networks is the vanishing and/or\nexploding gradient problem which is more prominently seen in training of\nRecurrent Neural Networks (RNNs). Thus several algorithms have been proposed\nfor training RNNs. This paper proposes a novel adaptive stochastic Nesterov\naccelerated quasiNewton (aSNAQ) method for training RNNs. The proposed method\naSNAQ is an accelerated method that uses the Nesterov's gradient term along\nwith second order curvature information. The performance of the proposed method\nis evaluated in Tensorflow on benchmark sequence modeling problems. The results\nshow an improved performance while maintaining a low per-iteration cost and\nthus can be effectively used to train RNNs.\n