Only a handful of the world's languages are abundant with the resources that\nenable practical applications of speech processing technologies. One of the\nmethods to overcome this problem is to use the resources existing in other\nlanguages to train a multilingual automatic speech recognition (ASR) model,\nwhich, intuitively, should learn some universal phonetic representations. In\nthis work, we focus on gaining a deeper understanding of how general these\nrepresentations might be, and how individual phones are getting improved in a\nmultilingual setting. To that end, we select a phonetically diverse set of\nlanguages, and perform a series of monolingual, multilingual and crosslingual\n(zero-shot) experiments. The ASR is trained to recognize the International\nPhonetic Alphabet (IPA) token sequences. We observe significant improvements\nacross all languages in the multilingual setting, and stark degradation in the\ncrosslingual setting, where the model, among other errors, considers Javanese\nas a tone language. Notably, as little as 10 hours of the target language\ntraining data tremendously reduces ASR error rates. Our analysis uncovered that\neven the phones that are unique to a single language can benefit greatly from\nadding training data from other languages - an encouraging result for the\nlow-resource speech community.\n
Paper
References (20)
Scroll for more · 8 remaining