Ultrasound-based Articulatory-to-Acoustic Mapping with WaveGlow Speech\n Synthesis

For articulatory-to-acoustic mapping using deep neural networks, typically\nspectral and excitation parameters of vocoders have been used as the training\ntargets. However, vocoding often results in buzzy and muffled final speech\nquality. Therefore, in this paper on ultrasound-based articulatory-to-acoustic\nconversion, we use a flow-based neural vocoder (WaveGlow) pre-trained on a\nlarge amount of English and Hungarian speech data. The inputs of the\nconvolutional neural network are ultrasound tongue images. The training target\nis the 80-dimensional mel-spectrogram, which results in a finer detailed\nspectral representation than the previously used 25-dimensional Mel-Generalized\nCepstrum. From the output of the ultrasound-to-mel-spectrogram prediction,\nWaveGlow inference results in synthesized speech. We compare the proposed\nWaveGlow-based system with a continuous vocoder which does not use strict\nvoiced/unvoiced decision when predicting F0. The results demonstrate that\nduring the articulatory-to-acoustic mapping experiments, the WaveGlow neural\nvocoder produces significantly more natural synthesized speech than the\nbaseline system. Besides, the advantage of WaveGlow is that F0 is included in\nthe mel-spectrogram representation, and it is not necessary to predict the\nexcitation separately.\n

Paper

References (34)

Scroll for more · 22 remaining

Similar papers

© 2026 NYSGPT2525 LLC