Don't Use English Dev: On the Zero-Shot Cross-Lingual Evaluation of Contextual Embeddings

Multilingual contextual embeddings have demonstrated state-of-the-art\nperformance in zero-shot cross-lingual transfer learning, where multilingual\nBERT is fine-tuned on one source language and evaluated on a different target\nlanguage. However, published results for mBERT zero-shot accuracy vary as much\nas 17 points on the MLDoc classification task across four papers. We show that\nthe standard practice of using English dev accuracy for model selection in the\nzero-shot setting makes it difficult to obtain reproducible results on the\nMLDoc and XNLI tasks. English dev accuracy is often uncorrelated (or even\nanti-correlated) with target language accuracy, and zero-shot performance\nvaries greatly at different points in the same fine-tuning run and between\ndifferent fine-tuning runs. These reproducibility issues are also present for\nother tasks with different pre-trained embeddings (e.g., MLQA with XLM-R). We\nrecommend providing oracle scores alongside zero-shot results: still fine-tune\nusing English data, but choose a checkpoint with the target dev set. Reporting\nthis upper bound makes results more consistent by avoiding arbitrarily bad\ncheckpoints.\n

Paper

References (29)

Scroll for more · 17 remaining

Similar papers

© 2026 NYSGPT2525 LLC