Investigating Deep Neural Structures and their Interpretability in the\n Domain of Voice Conversion
Generative Adversarial Networks (GANs) are machine learning networks based\naround creating synthetic data. Voice Conversion (VC) is a subset of voice\ntranslation that involves translating the paralinguistic features of a source\nspeaker to a target speaker while preserving the linguistic information. The\naim of non-parallel conditional GANs for VC is to translate an acoustic speech\nfeature sequence from one domain to another without the use of paired data. In\nthe study reported here, we investigated the interpretability of\nstate-of-the-art implementations of non-parallel GANs in the domain of VC. We\nshow that the learned representations in the repeating layers of a particular\nGAN architecture remain close to their original random initialised parameters,\ndemonstrating that it is the number of repeating layers that is more\nresponsible for the quality of the output. We also analysed the learned\nrepresentations of a model trained on one particular dataset when used during\ntransfer learning on another dataset. This showed extremely high levels of\nsimilarity across the entire network. Together, these results provide new\ninsight into how the learned representations of deep generative networks change\nduring learning and the importance in the number of layers.\n
Paper
References (49)
Scroll for more · 37 remaining