LNMap: Departures from Isomorphic Assumption in Bilingual Lexicon Induction Through Non-Linear Mapping in Latent Space
Most of the successful and predominant methods for bilingual lexicon\ninduction (BLI) are mapping-based, where a linear mapping function is learned\nwith the assumption that the word embedding spaces of different languages\nexhibit similar geometric structures (i.e., approximately isomorphic). However,\nseveral recent studies have criticized this simplified assumption showing that\nit does not hold in general even for closely related languages. In this work,\nwe propose a novel semi-supervised method to learn cross-lingual word\nembeddings for BLI. Our model is independent of the isomorphic assumption and\nuses nonlinear mapping in the latent space of two independently trained\nauto-encoders. Through extensive experiments on fifteen (15) different language\npairs (in both directions) comprising resource-rich and low-resource languages\nfrom two different datasets, we demonstrate that our method outperforms\nexisting models by a good margin. Ablation studies show the importance of\ndifferent model components and the necessity of non-linear mapping.\n
Paper
References (34)
Scroll for more · 22 remaining