Acoustic To Articulatory Speech Inversion Using Multi-Resolution Spectro-Temporal Representations Of Speech Signals
Multi-resolution spectro-temporal features of a speech signal represent how\nthe brain perceives sounds by tuning cortical cells to different spectral and\ntemporal modulations. These features produce a higher dimensional\nrepresentation of the speech signals. The purpose of this paper is to evaluate\nhow well the auditory cortex representation of speech signals contribute to\nestimate articulatory features of those corresponding signals. Since obtaining\narticulatory features from acoustic features of speech signals has been a\nchallenging topic of interest for different speech communities, we investigate\nthe possibility of using this multi-resolution representation of speech signals\nas acoustic features. We used U. of Wisconsin X-ray Microbeam (XRMB) database\nof clean speech signals to train a feed-forward deep neural network (DNN) to\nestimate articulatory trajectories of six tract variables. The optimal set of\nmulti-resolution spectro-temporal features to train the model were chosen using\nappropriate scale and rate vector parameters to obtain the best performing\nmodel. Experiments achieved a correlation of 0.675 with ground-truth tract\nvariables. We compared the performance of this speech inversion system with\nprior experiments conducted using Mel Frequency Cepstral Coefficients (MFCCs).\n
Paper
References (32)
Scroll for more · 20 remaining