Learning Cross-Lingual Sentence Representations via a Multi-task Dual-Encoder Model

A significant roadblock in multilingual neural language modeling is the lack\nof labeled non-English data. One potential method for overcoming this issue is\nlearning cross-lingual text representations that can be used to transfer the\nperformance from training on English tasks to non-English tasks, despite little\nto no task-specific non-English data. In this paper, we explore a natural setup\nfor learning cross-lingual sentence representations: the dual-encoder. We\nprovide a comprehensive evaluation of our cross-lingual representations on a\nnumber of monolingual, cross-lingual, and zero-shot/few-shot learning tasks,\nand also give an analysis of different learned cross-lingual embedding spaces.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC