Combining speakers of multiple languages to improve quality of neural voices

In this work, we explore multiple architectures and training procedures for\ndeveloping a multi-speaker and multi-lingual neural TTS system with the goals\nof a) improving the quality when the available data in the target language is\nlimited and b) enabling cross-lingual synthesis. We report results from a large\nexperiment using 30 speakers in 8 different languages across 15 different\nlocales. The system is trained on the same amount of data per speaker. Compared\nto a single-speaker model, when the suggested system is fine tuned to a\nspeaker, it produces significantly better quality in most of the cases while it\nonly uses less than $40\\%$ of the speaker's data used to build the\nsingle-speaker model. In cross-lingual synthesis, on average, the generated\nquality is within $80\\%$ of native single-speaker models, in terms of Mean\nOpinion Score.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC