Research on Acoustic Model of Speech Recognition Based on Neural Network with Improved Gating Unit
In the traditional speech recognition system, the speech acoustic model based on the recurrent neural network has limited ability to store long-distance historical information, and it is difficult to use the contextual relevance information of the speech. The standard long short-term memory has large scale, and the neural network training convergence speed is slow. To solve the above problem, this paper proposes a speech recognition acoustic model based on the bidirectional recurrent neural network with improved gated loop unit. Using the ReLU activation function instead of the hyperbolic tangent function, combined with the batch normalization method, helps the model to learn the long-term dependence of the network and maintain the stability of the output value. The appropriate network orthogonal initialization parameters further reduce the network training time and enhance the robustness of the acoustic model. Experimental results on the TIMIT and LibriSpeech datasets show that the improved gating recurrent unit model has a 2.8% absolute phoneme error rate reduction compared to the baseline system, compared to the standard long short-term memory model, the average iteration period of neural network training is reduced by 16.6%, which improves both recognition performance and computational efficiency.
Paper
Full text
Research on Acoustic Model of Speech Recognition Based on Neural Network with Improved Gating Unit
Semantic Scholar · Computer Science · 2019
Abstract
In the traditional speech recognition system, the speech acoustic model based on the recurrent neural network has limited ability to store long-distance historical information, and it is difficult to use the contextual relevance information of the speech. The standard long short-term memory has large scale, and the neural network training convergence speed is slow. To solve the above problem, this paper proposes a speech recognition acoustic model based on the bidirectional recurrent neural network with improved gated loop unit. Using the ReLU activation function instead of the hyperbolic tangent function, combined with the batch normalization method, helps the model to learn the long-term dependence of the network and maintain the stability of the output value. The appropriate network orthogonal initialization parameters further reduce the network training time and enhance the robustness of the acoustic model. Experimental results on the TIMIT and LibriSpeech datasets show that the improved gating recurrent unit model has a 2.8% absolute phoneme error rate reduction compared to the baseline system, compared to the standard long short-term memory model, the average iteration period of neural network training is reduced by 16.6%, which improves both recognition performance and computational efficiency.