Generating Multilingual Voices Using Speaker Space Translation Based on Bilingual Speaker Data

We present progress towards bilingual Text-to-Speech which is able to\ntransform a monolingual voice to speak a second language while preserving\nspeaker voice quality. We demonstrate that a bilingual speaker embedding space\ncontains a separate distribution for each language and that a simple transform\nin speaker space generated by the speaker embedding can be used to control the\ndegree of accent of a synthetic voice in a language. The same transform can be\napplied even to monolingual speakers.\n In our experiments speaker data from an English-Spanish (Mexican) bilingual\nspeaker was used, and the goal was to enable English speakers to speak Spanish\nand Spanish speakers to speak English. We found that the simple transform was\nsufficient to convert a voice from one language to the other with a high degree\nof naturalness. In one case the transformed voice outperformed a native\nlanguage voice in listening tests. Experiments further indicated that the\ntransform preserved many of the characteristics of the original voice. The\ndegree of accent present can be controlled and naturalness is relatively\nconsistent across a range of accent values.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC