In this work, we show the process of building a large-scale training set from\ndigital and digitized collections at a national library. The resulting\nBidirectional Encoder Representations from Transformers (BERT)-based language\nmodel for Norwegian outperforms multilingual BERT (mBERT) models in several\ntoken and sequence classification tasks for both Norwegian Bokm{\\aa}l and\nNorwegian Nynorsk. Our model also improves the mBERT performance for other\nlanguages present in the corpus such as English, Swedish, and Danish. For\nlanguages not included in the corpus, the weights degrade moderately while\nkeeping strong multilingual properties. Therefore, we show that building\nhigh-quality models within a memory institution using somewhat noisy optical\ncharacter recognition (OCR) content is feasible, and we hope to pave the way\nfor other memory institutions to follow.\n
Paper
References (44)
Scroll for more · 32 remaining