Cotatron: Transcription-Guided Speech Encoder for Any-to-Many Voice Conversion without Parallel Data

We propose Cotatron, a transcription-guided speech encoder for\nspeaker-independent linguistic representation. Cotatron is based on the\nmultispeaker TTS architecture and can be trained with conventional TTS\ndatasets. We train a voice conversion system to reconstruct speech with\nCotatron features, which is similar to the previous methods based on Phonetic\nPosteriorgram (PPG). By training and evaluating our system with 108 speakers\nfrom the VCTK dataset, we outperform the previous method in terms of both\nnaturalness and speaker similarity. Our system can also convert speech from\nspeakers that are unseen during training, and utilize ASR to automate the\ntranscription with minimal reduction of the performance. Audio samples are\navailable at https://mindslab-ai.github.io/cotatron, and the code with a\npre-trained model will be made available soon.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC