Improving Recognition for Disordered Speech in Indonesian Language: a Data Augmentation approach

For individuals with speech disorders, speech recognition serves as a crucial communication tool. But most speech recognition systems are trained on normal speech data, worsen by limited disordered speech data which mostly available in English language. Data augmentation is one possible approach, where we can apply signal processing such as speed perturbation to transform normal speech into disordered speech. In this work, we studied how augmented data composition in a dataset affects the system performance on recognizing disordered speech. Using QuartzNet CNN as the acoustic model, we evaluated the speech recognition system on an augmented dataset built based on an Indonesian language speech database from Mozilla Corpus. Using speed perturbation to generate disordered speech, the initial results show that augmenting normal speech dataset with 25-50% more disordered speech data could help improve the Word Error Rate (WER) of the model in recognizing disordered speech. This result, however, is limited as we use the same speed perturbation method for training and testing datasets.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC