2021/03/12 by Adnan Haider, Chao Zhang, Haider, Adnan +5
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2103.07554
openalex publication_date 2021/03/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper presents a novel natural gradient and Hessian-free (NGHF)\noptimisation framework for neural network training that can operate efficiently\nin a distributed manner. It relies on the linear conjugate gradient (CG)\nalgorithm to combine the natural gradient (NG) method with local curvature\ninformation from Hessian-free (HF) or other second-order methods. A solution to\na numerical issue in CG allows effective parameter updates to be generated with\nfar fewer CG iterations than usually used (e.g. 5-8 instead of 200). This work\nalso presents a novel preconditioning approach to improve the progress made by\nindividual CG iterations for models with shared parameters. Although applicable\nto other training losses and model structures, NGHF is investigated in this\npaper for lattice-based discriminative sequence training for hybrid hidden\nMarkov model acoustic models using a standard recurrent neural network, long\nshort-term memory, and time delay neural network models for output probability\ncalculation. Automatic speech recognition experiments are reported on the\nmulti-genre broadcast data set for a range of different acoustic model types.\nThese experiments show that NGHF achieves larger word error rate reductions\nthan standard stochastic gradient descent or Adam, while requiring orders of\nmagnitude fewer parameter updates.\n