Speaker Conditional WaveRNN: Towards Universal Neural Vocoder for Unseen Speaker and Recording Conditions

Recent advancements in deep learning led to human-level performance in\nsingle-speaker speech synthesis. However, there are still limitations in terms\nof speech quality when generalizing those systems into multiple-speaker models\nespecially for unseen speakers and unseen recording qualities. For instance,\nconventional neural vocoders are adjusted to the training speaker and have poor\ngeneralization capabilities to unseen speakers. In this work, we propose a\nvariant of WaveRNN, referred to as speaker conditional WaveRNN (SC-WaveRNN). We\ntarget towards the development of an efficient universal vocoder even for\nunseen speakers and recording conditions. In contrast to standard WaveRNN,\nSC-WaveRNN exploits additional information given in the form of speaker\nembeddings. Using publicly-available data for training, SC-WaveRNN achieves\nsignificantly better performance over baseline WaveRNN on both subjective and\nobjective metrics. In MOS, SC-WaveRNN achieves an improvement of about 23% for\nseen speaker and seen recording condition and up to 95% for unseen speaker and\nunseen condition. Finally, we extend our work by implementing a multi-speaker\ntext-to-speech (TTS) synthesis similar to zero-shot speaker adaptation. In\nterms of performance, our system has been preferred over the baseline TTS\nsystem by 60% over 15.5% and by 60.9% over 32.6%, for seen and unseen speakers,\nrespectively.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC