Ultra2Speech -- A Deep Learning Framework for Formant Frequency Estimation and Tracking from Ultrasound Tongue Images
Thousands of individuals need surgical removal of their larynx due to\ncritical diseases every year and therefore, require an alternative form of\ncommunication to articulate speech sounds after the loss of their voice box.\nThis work addresses the articulatory-to-acoustic mapping problem based on\nultrasound (US) tongue images for the development of a silent-speech interface\n(SSI) that can provide them with an assistance in their daily interactions. Our\napproach targets automatically extracting tongue movement information by\nselecting an optimal feature set from US images and mapping these features to\nthe acoustic space. We use a novel deep learning architecture to map US tongue\nimages from the US probe placed beneath a subject's chin to formants that we\ncall, Ultrasound2Formant (U2F) Net. It uses hybrid spatio-temporal 3D\nconvolutions followed by feature shuffling, for the estimation and tracking of\nvowel formants from US images. The formant values are then utilized to\nsynthesize continuous time-varying vowel trajectories, via Klatt Synthesizer.\nOur best model achieves R-squared (R^2) measure of 99.96% for the regression\ntask. Our network lays the foundation for an SSI as it successfully tracks the\ntongue contour automatically as an internal representation without any explicit\nannotation.\n