Significance of Pitch-Based Spectral Normalization for Children's Speech Recognition

It is well known from the literature that due to several acoustic mismatches, the recognition performances of children's speech using adult-trained-acoustic models get deteriorated. The differences in pitch and speaking rate are the two major factors that cause the acoustic mismatch between two groups of speakers. This work proposes to incorporate pitch information into an automatic speech recognition (ASR) system by exploiting the correlation between pitch and formants. By using the pitch-based spectrum normalization module in the front-end feature extraction process, the performance of mismatch ASR system is improved for children of different age groups. Further, fuzzy-based time scale modification is applied to study the effect of speaking-rate normalization on the proposed feature. The proposed feature results in relative improvement of <inline-formula><tex-math notation="LaTeX">$\text{30}{\%}$</tex-math></inline-formula> and <inline-formula><tex-math notation="LaTeX">$\text{33}{\%}$</tex-math></inline-formula> on DLSTM-based ASR system over the MFCC baseline without and with speaking-rate normalization, respectively.

Paper

Full text

PDF

Significance of Pitch-Based Spectral Normalization for Children's Speech Recognition

Semantic Scholar · Computer Science · 2019

Abstract

It is well known from the literature that due to several acoustic mismatches, the recognition performances of children's speech using adult-trained-acoustic models get deteriorated. The differences in pitch and speaking rate are the two major factors that cause the acoustic mismatch between two groups of speakers. This work proposes to incorporate pitch information into an automatic speech recognition (ASR) system by exploiting the correlation between pitch and formants. By using the pitch-based spectrum normalization module in the front-end feature extraction process, the performance of mismatch ASR system is improved for children of different age groups. Further, fuzzy-based time scale modification is applied to study the effect of speaking-rate normalization on the proposed feature. The proposed feature results in relative improvement of <inline-formula><tex-math notation="LaTeX">$\text{30}{%}$</tex-math></inline-formula> and <inline-formula><tex-math notation="LaTeX">$\text{33}{%}$</tex-math></inline-formula> on DLSTM-based ASR system over the MFCC baseline without and with speaking-rate normalization, respectively.

Similar papers

© 2026 NYSGPT2525 LLC