Multilingual Machine Translation: Closing the Gap between Shared and Language-specific Encoder-Decoders
State-of-the-art multilingual machine translation relies on a universal\nencoder-decoder, which requires retraining the entire system to add new\nlanguages. In this paper, we propose an alternative approach that is based on\nlanguage-specific encoder-decoders, and can thus be more easily extended to new\nlanguages by learning their corresponding modules. So as to encourage a common\ninterlingua representation, we simultaneously train the N initial languages.\nOur experiments show that the proposed approach outperforms the universal\nencoder-decoder by 3.28 BLEU points on average, and when adding new languages,\nwithout the need to retrain the rest of the modules. All in all, our work\ncloses the gap between shared and language-specific encoder-decoders, advancing\ntoward modular multilingual machine translation systems that can be flexibly\nextended in lifelong learning settings.\n