WikiLingua: A New Benchmark Dataset for Cross-Lingual Abstractive Summarization

We introduce WikiLingua, a large-scale, multilingual dataset for the\nevaluation of crosslingual abstractive summarization systems. We extract\narticle and summary pairs in 18 languages from WikiHow, a high quality,\ncollaborative resource of how-to guides on a diverse set of topics written by\nhuman authors. We create gold-standard article-summary alignments across\nlanguages by aligning the images that are used to describe each how-to step in\nan article. As a set of baselines for further studies, we evaluate the\nperformance of existing cross-lingual abstractive summarization methods on our\ndataset. We further propose a method for direct crosslingual summarization\n(i.e., without requiring translation at inference time) by leveraging synthetic\ndata and Neural Machine Translation as a pre-training step. Our method\nsignificantly outperforms the baseline approaches, while being more cost\nefficient during inference.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC