Learning robust speech representation with an articulatory-regularized variational autoencoder
It is increasingly considered that human speech perception and production\nboth rely on articulatory representations. In this paper, we investigate\nwhether this type of representation could improve the performances of a deep\ngenerative model (here a variational autoencoder) trained to encode and decode\nacoustic speech features. First we develop an articulatory model able to\nassociate articulatory parameters describing the jaw, tongue, lips and velum\nconfigurations with vocal tract shapes and spectral features. Then we\nincorporate these articulatory parameters into a variational autoencoder\napplied on spectral features by using a regularization technique that\nconstraints part of the latent space to follow articulatory trajectories. We\nshow that this articulatory constraint improves model training by decreasing\ntime to convergence and reconstruction loss at convergence, and yields better\nperformance in a speech denoising task.\n