Bridging Linguistic Typology and Multilingual Machine Translation with Multi-View Language Representations
Sparse language vectors from linguistic typology databases and learned\nembeddings from tasks like multilingual machine translation have been\ninvestigated in isolation, without analysing how they could benefit from each\nother's language characterisation. We propose to fuse both views using singular\nvector canonical correlation analysis and study what kind of information is\ninduced from each source. By inferring typological features and language\nphylogenies, we observe that our representations embed typology and strengthen\ncorrelations with language relationships. We then take advantage of our\nmulti-view language vector space for multilingual machine translation, where we\nachieve competitive overall translation accuracy in tasks that require\ninformation about language similarities, such as language clustering and\nranking candidates for multilingual transfer. With our method, which is also\nreleased as a tool, we can easily project and assess new languages without\nexpensive retraining of massive multilingual or ranking models, which are major\ndisadvantages of related approaches.\n