A Distributed Optimisation Framework Combining Natural Gradient with Hessian-Free for Discriminative Sequence Training
This paper presents a novel natural gradient and Hessian-free (NGHF)\noptimisation framework for neural network training that can operate efficiently\nin a distributed manner. It relies on the linear conjugate gradient (CG)\nalgorithm to combine the natural gradient (NG) method with local curvature\ninformation from Hessian-free (HF) or other second-order methods. A solution to\na numerical issue in CG allows effective parameter updates to be generated with\nfar fewer CG iterations than usually used (e.g. 5-8 instead of 200). This work\nalso presents a novel preconditioning approach to improve the progress made by\nindividual CG iterations for models with shared parameters. Although applicable\nto other training losses and model structures, NGHF is investigated in this\npaper for lattice-based discriminative sequence training for hybrid hidden\nMarkov model acoustic models using a standard recurrent neural network, long\nshort-term memory, and time delay neural network models for output probability\ncalculation. Automatic speech recognition experiments are reported on the\nmulti-genre broadcast data set for a range of different acoustic model types.\nThese experiments show that NGHF achieves larger word error rate reductions\nthan standard stochastic gradient descent or Adam, while requiring orders of\nmagnitude fewer parameter updates.\n
Paper
References (99)
Scroll for more · 38 remaining