Enabling Zero-shot Multilingual Spoken Language Translation with Language-Specific Encoders and Decoders
Current end-to-end approaches to Spoken Language Translation (SLT) rely on\nlimited training resources, especially for multilingual settings. On the other\nhand, Multilingual Neural Machine Translation (MultiNMT) approaches rely on\nhigher-quality and more massive data sets. Our proposed method extends a\nMultiNMT architecture based on language-specific encoders-decoders to the task\nof Multilingual SLT (MultiSLT). Our method entirely eliminates the dependency\nfrom MultiSLT data and it is able to translate while training only on ASR and\nMultiNMT data.\n Our experiments on four different languages show that coupling the speech\nencoder to the MultiNMT architecture produces similar quality translations\ncompared to a bilingual baseline ($\\pm 0.2$ BLEU) while effectively allowing\nfor zero-shot MultiSLT. Additionally, we propose using an Adapter module for\ncoupling the speech inputs. This Adapter module produces consistent\nimprovements up to +6 BLEU points on the proposed architecture and +1 BLEU\npoint on the end-to-end baseline.\n
Paper
References (23)
Scroll for more · 11 remaining