Evaluating Multilingual Text Encoders for Unsupervised Cross-Lingual Retrieval

Pretrained multilingual text encoders based on neural Transformer\narchitectures, such as multilingual BERT (mBERT) and XLM, have achieved strong\nperformance on a myriad of language understanding tasks. Consequently, they\nhave been adopted as a go-to paradigm for multilingual and cross-lingual\nrepresentation learning and transfer, rendering cross-lingual word embeddings\n(CLWEs) effectively obsolete. However, questions remain to which extent this\nfinding generalizes 1) to unsupervised settings and 2) for ad-hoc cross-lingual\nIR (CLIR) tasks. Therefore, in this work we present a systematic empirical\nstudy focused on the suitability of the state-of-the-art multilingual encoders\nfor cross-lingual document and sentence retrieval tasks across a large number\nof language pairs. In contrast to supervised language understanding, our\nresults indicate that for unsupervised document-level CLIR -- a setup with no\nrelevance judgments for IR-specific fine-tuning -- pretrained encoders fail to\nsignificantly outperform models based on CLWEs. For sentence-level CLIR, we\ndemonstrate that state-of-the-art performance can be achieved. However, the\npeak performance is not met using the general-purpose multilingual text\nencoders `off-the-shelf', but rather relying on their variants that have been\nfurther specialized for sentence understanding tasks.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC