FragmentVC: Any-to-Any Voice Conversion by End-to-End Extracting and Fusing Fine-Grained Voice Fragments With Attention
Any-to-any voice conversion aims to convert the voice from and to any\nspeakers even unseen during training, which is much more challenging compared\nto one-to-one or many-to-many tasks, but much more attractive in real-world\nscenarios. In this paper we proposed FragmentVC, in which the latent phonetic\nstructure of the utterance from the source speaker is obtained from Wav2Vec\n2.0, while the spectral features of the utterance(s) from the target speaker\nare obtained from log mel-spectrograms. By aligning the hidden structures of\nthe two different feature spaces with a two-stage training process, FragmentVC\nis able to extract fine-grained voice fragments from the target speaker\nutterance(s) and fuse them into the desired utterance, all based on the\nattention mechanism of Transformer as verified with analysis on attention maps,\nand is accomplished end-to-end. This approach is trained with reconstruction\nloss only without any disentanglement considerations between content and\nspeaker information and doesn't require parallel data. Objective evaluation\nbased on speaker verification and subjective evaluation with MOS both showed\nthat this approach outperformed SOTA approaches, such as AdaIN-VC and AutoVC.\n
Paper
References (28)
Scroll for more · 16 remaining