Research on Acoustic Model of Speech Recognition Based on Neural Network with Improved Gating Unit

In the traditional speech recognition system, the speech acoustic model based on the recurrent neural network has limited ability to store long-distance historical information, and it is difficult to use the contextual relevance information of the speech. The standard long short-term memory has large scale, and the neural network training convergence speed is slow. To solve the above problem, this paper proposes a speech recognition acoustic model based on the bidirectional recurrent neural network with improved gated loop unit. Using the ReLU activation function instead of the hyperbolic tangent function, combined with the batch normalization method, helps the model to learn the long-term dependence of the network and maintain the stability of the output value. The appropriate network orthogonal initialization parameters further reduce the network training time and enhance the robustness of the acoustic model. Experimental results on the TIMIT and LibriSpeech datasets show that the improved gating recurrent unit model has a 2.8% absolute phoneme error rate reduction compared to the baseline system, compared to the standard long short-term memory model, the average iteration period of neural network training is reduced by 16.6%, which improves both recognition performance and computational efficiency.

Paper

Full text

PDF

Research on Acoustic Model of Speech Recognition Based on Neural Network with Improved Gating Unit

Semantic Scholar · Computer Science · 2019

Abstract

In the traditional speech recognition system, the speech acoustic model based on the recurrent neural network has limited ability to store long-distance historical information, and it is difficult to use the contextual relevance information of the speech. The standard long short-term memory has large scale, and the neural network training convergence speed is slow. To solve the above problem, this paper proposes a speech recognition acoustic model based on the bidirectional recurrent neural network with improved gated loop unit. Using the ReLU activation function instead of the hyperbolic tangent function, combined with the batch normalization method, helps the model to learn the long-term dependence of the network and maintain the stability of the output value. The appropriate network orthogonal initialization parameters further reduce the network training time and enhance the robustness of the acoustic model. Experimental results on the TIMIT and LibriSpeech datasets show that the improved gating recurrent unit model has a 2.8% absolute phoneme error rate reduction compared to the baseline system, compared to the standard long short-term memory model, the average iteration period of neural network training is reduced by 16.6%, which improves both recognition performance and computational efficiency.

Similar papers

© 2026 NYSGPT2525 LLC