Transfer Learning from Speech Synthesis to Voice Conversion with\n Non-Parallel Training Data
This paper presents a novel framework to build a voice conversion (VC) system\nby learning from a text-to-speech (TTS) synthesis system, that is called TTS-VC\ntransfer learning. We first develop a multi-speaker speech synthesis system\nwith sequence-to-sequence encoder-decoder architecture, where the encoder\nextracts robust linguistic representations of text, and the decoder,\nconditioned on target speaker embedding, takes the context vectors and the\nattention recurrent network cell output to generate target acoustic features.\nWe take advantage of the fact that TTS system maps input text to speaker\nindependent context vectors, and reuse such a mapping to supervise the training\nof latent representations of an encoder-decoder voice conversion system. In the\nvoice conversion system, the encoder takes speech instead of text as input,\nwhile the decoder is functionally similar to TTS decoder. As we condition the\ndecoder on speaker embedding, the system can be trained on non-parallel data\nfor any-to-any voice conversion. During voice conversion training, we present\nboth text and speech to speech synthesis and voice conversion networks\nrespectively. At run-time, the voice conversion network uses its own\nencoder-decoder architecture. Experiments show that the proposed approach\noutperforms two competitive voice conversion baselines consistently, namely\nphonetic posteriorgram and variational autoencoder methods, in terms of speech\nquality, naturalness, and speaker similarity.\n