Improving Low Resource Code-switched ASR using Augmented Code-switched TTS

Building Automatic Speech Recognition (ASR) systems for code-switched speech\nhas recently gained renewed attention due to the widespread use of speech\ntechnologies in multilingual communities worldwide. End-to-end ASR systems are\na natural modeling choice due to their ease of use and superior performance in\nmonolingual settings. However, it is well known that end-to-end systems require\nlarge amounts of labeled speech. In this work, we investigate improving\ncode-switched ASR in low resource settings via data augmentation using\ncode-switched text-to-speech (TTS) synthesis. We propose two targeted\ntechniques to effectively leverage TTS speech samples: 1) Mixup, an existing\ntechnique to create new training samples via linear interpolation of existing\nsamples, applied to TTS and real speech samples, and 2) a new loss function,\nused in conjunction with TTS samples, to encourage code-switched predictions.\nWe report significant improvements in ASR performance achieving absolute word\nerror rate (WER) reductions of up to 5%, and measurable improvement in code\nswitching using our proposed techniques on a Hindi-English code-switched ASR\ntask.\n

Paper

References (38)

Scroll for more · 26 remaining

Similar papers

© 2026 NYSGPT2525 LLC