Investigating on Incorporating Pretrained and Learnable Speaker Representations for Multi-Speaker Multi-Style Text-to-Speech

The few-shot multi-speaker multi-style voice cloning task is to synthesize\nutterances with voice and speaking style similar to a reference speaker given\nonly a few reference samples. In this work, we investigate different speaker\nrepresentations and proposed to integrate pretrained and learnable speaker\nrepresentations. Among different types of embeddings, the embedding pretrained\nby voice conversion achieves the best performance. The FastSpeech 2 model\ncombined with both pretrained and learnable speaker representations shows great\ngeneralization ability on few-shot speakers and achieved 2nd place in the\none-shot track of the ICASSP 2021 M2VoC challenge.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC