A common approach to the automatic detection of mispronunciation in language\nlearning is to recognize the phonemes produced by a student and compare it to\nthe expected pronunciation of a native speaker. This approach makes two\nsimplifying assumptions: a) phonemes can be recognized from speech with high\naccuracy, b) there is a single correct way for a sentence to be pronounced.\nThese assumptions do not always hold, which can result in a significant amount\nof false mispronunciation alarms. We propose a novel approach to overcome this\nproblem based on two principles: a) taking into account uncertainty in the\nautomatic phoneme recognition step, b) accounting for the fact that there may\nbe multiple valid pronunciations. We evaluate the model on non-native (L2)\nEnglish speech of German, Italian and Polish speakers, where it is shown to\nincrease the precision of detecting mispronunciations by up to 18% (relative)\ncompared to the common approach.\n
Paper
References (21)
Scroll for more · 9 remaining