Quantization of Acoustic Model Parameters in Automatic Speech\n Recognition Framework

State-of-the-art hybrid automatic speech recognition (ASR) system exploits\ndeep neural network (DNN) based acoustic models (AM) trained with Lattice\nFree-Maximum Mutual Information (LF-MMI) criterion and n-gram language models.\nThe AMs typically have millions of parameters and require significant parameter\nreduction to operate on embedded devices. The impact of parameter quantization\non the overall word recognition performance is studied in this paper. Following\napproaches are presented: (i) AM trained in Kaldi framework with conventional\nfactorized TDNN (TDNN-F) architecture, (ii) the TDNN AM built in Kaldi loaded\ninto the PyTorch toolkit using a C++ wrapper for post-training quantization,\n(iii) quantization-aware training in PyTorch for Kaldi TDNN model, (iv)\nquantization-aware training in Kaldi. Results obtained on standard Librispeech\nsetup provide an interesting overview of recognition accuracy w.r.t. applied\nquantization scheme.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC