A comparable study of modeling units for end-to-end Mandarin speech recognition

End-to-end speech recognition has become increasingly popular in Mandarin speech recognition and achieved delightful performance. Mandarin is a tonal language which is different from English and requires special treatment for the acoustic modeling units. There have been several different kinds of modeling units for Mandarin such as phoneme, syllable and Chinese character. In this work, we explore two major end-to-end models: connectionist temporal classification (CTC) model and attention based encoder-decoder model for Mandarin speech recognition. We compare the performance of three different scaled modeling units: context dependent phoneme (CDP), syllable with tone and Chinese character. We find that all types of modeling units can achieve similar character error rate (CER) in CTC model and the performance of Chinese character attention model is better than syllable attention model. Furthermore, we find that Chinese character is a reasonable unit for Mandarin speech recognition. On HKUST task, Chinese character attention model achieves a CER of 35.2% and CTC model gets a CER of 35.7%, on internal Callcenter task, CERs are 5.68% and 7.29%, respectively, on internal Reading task, CERs are 4.89% and 5.79%, respectively. Moreover, attention model achieves a better performance than CTC model on these datasets.

Paper

References (30)

Scroll for more · 18 remaining

Similar papers

© 2026 NYSGPT2525 LLC