The Tatoeba Translation Challenge -- Realistic Data Sets for Low Resource and Multilingual MT
This paper describes the development of a new benchmark for machine\ntranslation that provides training and test data for thousands of language\npairs covering over 500 languages and tools for creating state-of-the-art\ntranslation models from that collection. The main goal is to trigger the\ndevelopment of open translation tools and models with a much broader coverage\nof the World's languages. Using the package it is possible to work on realistic\nlow-resource scenarios avoiding artificially reduced setups that are common\nwhen demonstrating zero-shot or few-shot learning. For the first time, this\npackage provides a comprehensive collection of diverse data sets in hundreds of\nlanguages with systematic language and script annotation and data splits to\nextend the narrow coverage of existing benchmarks. Together with the data\nrelease, we also provide a growing number of pre-trained baseline models for\nindividual language pairs and selected language groups.\n