Operationalizing a National Digital Library: The Case for a Norwegian Transformer Model

In this work, we show the process of building a large-scale training set from\ndigital and digitized collections at a national library. The resulting\nBidirectional Encoder Representations from Transformers (BERT)-based language\nmodel for Norwegian outperforms multilingual BERT (mBERT) models in several\ntoken and sequence classification tasks for both Norwegian Bokm{\\aa}l and\nNorwegian Nynorsk. Our model also improves the mBERT performance for other\nlanguages present in the corpus such as English, Swedish, and Danish. For\nlanguages not included in the corpus, the weights degrade moderately while\nkeeping strong multilingual properties. Therefore, we show that building\nhigh-quality models within a memory institution using somewhat noisy optical\ncharacter recognition (OCR) content is feasible, and we hope to pave the way\nfor other memory institutions to follow.\n

Paper

References (44)

Scroll for more · 32 remaining

Similar papers

© 2026 NYSGPT2525 LLC