One Model to Pronounce Them All: Multilingual Grapheme-to-Phoneme Conversion With a Transformer Ensemble
The task of grapheme-to-phoneme (G2P) conversion is important for both speech\nrecognition and synthesis. Similar to other speech and language processing\ntasks, in a scenario where only small-sized training data are available,\nlearning G2P models is challenging. We describe a simple approach of exploiting\nmodel ensembles, based on multilingual Transformers and self-training, to\ndevelop a highly effective G2P solution for 15 languages. Our models are\ndeveloped as part of our participation in the SIGMORPHON 2020 Shared Task 1\nfocused at G2P. Our best models achieve 14.99 word error rate (WER) and 3.30\nphoneme error rate (PER), a sizeable improvement over the shared task\ncompetitive baselines.\n