Adapt-and-Adjust: Overcoming the Long-Tail Problem of Multilingual Speech Recognition

One crucial challenge of real-world multilingual speech recognition is the\nlong-tailed distribution problem, where some resource-rich languages like\nEnglish have abundant training data, but a long tail of low-resource languages\nhave varying amounts of limited training data. To overcome the long-tail\nproblem, in this paper, we propose Adapt-and-Adjust (A2), a transformer-based\nmulti-task learning framework for end-to-end multilingual speech recognition.\nThe A2 framework overcomes the long-tail problem via three techniques: (1)\nexploiting a pretrained multilingual language model (mBERT) to improve the\nperformance of low-resource languages; (2) proposing dual adapters consisting\nof both language-specific and language-agnostic adaptation with minimal\nadditional parameters; and (3) overcoming the class imbalance, either by\nimposing class priors in the loss during training or adjusting the logits of\nthe softmax output during inference. Extensive experiments on the CommonVoice\ncorpus show that A2 significantly outperforms conventional approaches.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC