Efficient neural speech synthesis for low-resource languages through multilingual modeling

Recent advances in neural TTS have led to models that can produce\nhigh-quality synthetic speech. However, these models typically require large\namounts of training data, which can make it costly to produce a new voice with\nthe desired quality. Although multi-speaker modeling can reduce the data\nrequirements necessary for a new voice, this approach is usually not viable for\nmany low-resource languages for which abundant multi-speaker data is not\navailable. In this paper, we therefore investigated to what extent multilingual\nmulti-speaker modeling can be an alternative to monolingual multi-speaker\nmodeling, and explored how data from foreign languages may best be combined\nwith low-resource language data. We found that multilingual modeling can\nincrease the naturalness of low-resource language speech, showed that\nmultilingual models can produce speech with a naturalness comparable to\nmonolingual multi-speaker models, and saw that the target language naturalness\nwas affected by the strategy used to add foreign language data.\n

Paper

References (24)

Scroll for more · 12 remaining

Similar papers

© 2026 NYSGPT2525 LLC