WaveTract: A hybrid generative model for speech synthesis

In this paper we propose a hybrid speech synthesizer, which combines a neural network based synthesizer with an acoustical vocal tract model. Low frequency audio of the speech is generated with a WaveGlow. This audio is then upsampled by a simulated vocal tract. The proposed hybrid model can be trained as any supervised neural network. Since the WaveGlow used in this proposed model is smaller and the vocal tract model can be implemented efficiently, it produces audio at a much higher speed and delivers a decent audio quality. The proposed model also makes it possible to divide the inference into two parts. The upsampling process can be done on a low performance CPU at the client side, for example mobile phones or home assistant devices. This way the required performance at server side will also be lower.

Paper

Full text

PDF

WaveTract: A hybrid generative model for speech synthesis

Semantic Scholar · Computer Science · 2019

Abstract

In this paper we propose a hybrid speech synthesizer, which combines a neural network based synthesizer with an acoustical vocal tract model. Low frequency audio of the speech is generated with a WaveGlow. This audio is then upsampled by a simulated vocal tract. The proposed hybrid model can be trained as any supervised neural network. Since the WaveGlow used in this proposed model is smaller and the vocal tract model can be implemented efficiently, it produces audio at a much higher speed and delivers a decent audio quality. The proposed model also makes it possible to divide the inference into two parts. The upsampling process can be done on a low performance CPU at the client side, for example mobile phones or home assistant devices. This way the required performance at server side will also be lower.

Similar papers

© 2026 NYSGPT2525 LLC