Multi-stage Multi-modalities Fusion of Lip, Tongue and Acoustics Information for Speech Recognition
The ultrasound tongue imaging (UTI) and lip video are commonly used to capture the acoustic clue to obtain the visual articulatory information of speakers. However, single signal of UTI or lip video cannot completely represent the pronunciation process of speakers. In this paper, we proposed to use the convolutional neural network (CNN)-based framework to fuse the lip and tongue movement information to represent the pronunciation process of speakers. In addition, we designed the multi-stage fusion framework (MF-SR) to fuse the lip-tongue visual information and the acoustic features extracted from the speech. To evaluate our proposed method, we designed the data stream comparative experiments, the speech pattern comparative experiments, and the data increment experiments based on TAL1 dataset. The results show that the best word error rate (WER) of our proposed method on audio-visual speech recognition task is 20.03%. The best WER of our proposed model on visual-only speech recognition task is 23.34%, which is reduced by 1.75% compared with the baseline method. The results illustrate that our proposed method can effectively further improve the performance of the lip-tongue-audio fusion speech recognition.
Paper
Full text
Multi-stage Multi-modalities Fusion of Lip, Tongue and Acoustics Information for Speech Recognition
Semantic Scholar · Computer Science · 2023
Abstract
The ultrasound tongue imaging (UTI) and lip video are commonly used to capture the acoustic clue to obtain the visual articulatory information of speakers. However, single signal of UTI or lip video cannot completely represent the pronunciation process of speakers. In this paper, we proposed to use the convolutional neural network (CNN)-based framework to fuse the lip and tongue movement information to represent the pronunciation process of speakers. In addition, we designed the multi-stage fusion framework (MF-SR) to fuse the lip-tongue visual information and the acoustic features extracted from the speech. To evaluate our proposed method, we designed the data stream comparative experiments, the speech pattern comparative experiments, and the data increment experiments based on TAL1 dataset. The results show that the best word error rate (WER) of our proposed method on audio-visual speech recognition task is 20.03%. The best WER of our proposed model on visual-only speech recognition task is 23.34%, which is reduced by 1.75% compared with the baseline method. The results illustrate that our proposed method can effectively further improve the performance of the lip-tongue-audio fusion speech recognition.
References (41)
Scroll for more · 29 remaining