One Question Answering Model for Many Languages with Cross-lingual Dense Passage Retrieval

We present Cross-lingual Open-Retrieval Answer Generation (CORA), the first\nunified many-to-many question answering (QA) model that can answer questions\nacross many languages, even for ones without language-specific annotated data\nor knowledge sources. We introduce a new dense passage retrieval algorithm that\nis trained to retrieve documents across languages for a question. Combined with\na multilingual autoregressive generation model, CORA answers directly in the\ntarget language without any translation or in-language retrieval modules as\nused in prior work. We propose an iterative training method that automatically\nextends annotated data available only in high-resource languages to\nlow-resource ones. Our results show that CORA substantially outperforms the\nprevious state of the art on multilingual open QA benchmarks across 26\nlanguages, 9 of which are unseen during training. Our analyses show the\nsignificance of cross-lingual retrieval and generation in many languages,\nparticularly under low-resource settings.\n

Paper

References (61)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC