End-to-End Speaker Height and age estimation using Attention Mechanism with LSTM-RNN

Automatic height and age estimation of speakers using acoustic features is\nwidely used for the purpose of human-computer interaction, forensics, etc. In\nthis work, we propose a novel approach of using attention mechanism to build an\nend-to-end architecture for height and age estimation. The attention mechanism\nis combined with Long Short-Term Memory(LSTM) encoder which is able to capture\nlong-term dependencies in the input acoustic features. We modify the\nconventionally used Attention -- which calculates context vectors the sum of\nattention only across timeframes -- by introducing a modified context vector\nwhich takes into account total attention across encoder units as well, giving\nus a new cross-attention mechanism. Apart from this, we also investigate a\nmulti-task learning approach for jointly estimating speaker height and age. We\ntrain and test our model on the TIMIT corpus. Our model outperforms several\napproaches in the literature. We achieve a root mean square error (RMSE) of\n6.92cm and6.34cm for male and female heights respectively and RMSE of 7.85years\nand 8.75years for male and females ages respectively. By tracking the attention\nweights allocated to different phones, we find that Vowel phones are most\nimportant whistlestop phones are least important for the estimation task.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC