Reading comprehension is a well studied task, with huge training datasets in\nEnglish. This work focuses on building reading comprehension systems for Czech,\nwithout requiring any manually annotated Czech training data. First of all, we\nautomatically translated SQuAD 1.1 and SQuAD 2.0 datasets to Czech to create\ntraining and development data, which we release at\nhttp://hdl.handle.net/11234/1-3249. We then trained and evaluated several BERT\nand XLM-RoBERTa baseline models. However, our main focus lies in cross-lingual\ntransfer models. We report that a XLM-RoBERTa model trained on English data and\nevaluated on Czech achieves very competitive performance, only approximately 2\npercent points worse than a~model trained on the translated Czech data. This\nresult is extremely good, considering the fact that the model has not seen any\nCzech data during training. The cross-lingual transfer approach is very\nflexible and provides a reading comprehension in any language, for which we\nhave enough monolingual raw texts.\n
Paper
References (15)
Scroll for more · 3 remaining