Cross-Domain Adaptation of Spoken Language Identification for Related Languages: The Curious Case of Slavic Languages
State-of-the-art spoken language identification (LID) systems, which are\nbased on end-to-end deep neural networks, have shown remarkable success not\nonly in discriminating between distant languages but also between\nclosely-related languages or even different spoken varieties of the same\nlanguage. However, it is still unclear to what extent neural LID models\ngeneralize to speech samples with different acoustic conditions due to domain\nshift. In this paper, we present a set of experiments to investigate the impact\nof domain mismatch on the performance of neural LID systems for a subset of six\nSlavic languages across two domains (read speech and radio broadcast) and\nexamine two low-level signal descriptors (spectral and cepstral features) for\nthis task. Our experiments show that (1) out-of-domain speech samples severely\nhinder the performance of neural LID models, and (2) while both spectral and\ncepstral features show comparable performance within-domain, spectral features\nshow more robustness under domain mismatch. Moreover, we apply unsupervised\ndomain adaptation to minimize the discrepancy between the two domains in our\nstudy. We achieve relative accuracy improvements that range from 9% to 77%\ndepending on the diversity of acoustic conditions in the source domain.\n
Paper
References (34)
Scroll for more · 22 remaining