ÚFAL at MultiLexNorm 2021: Improving Multilingual Lexical Normalization by Fine-tuning ByT5

We present the winning entry to the Multilingual Lexical Normalization\n(MultiLexNorm) shared task at W-NUT 2021 (van der Goot et al., 2021a), which\nevaluates lexical-normalization systems on 12 social media datasets in 11\nlanguages. We base our solution on a pre-trained byte-level language model,\nByT5 (Xue et al., 2021a), which we further pre-train on synthetic data and then\nfine-tune on authentic normalization data. Our system achieves the best\nperformance by a wide margin in intrinsic evaluation, and also the best\nperformance in extrinsic evaluation through dependency parsing. The source code\nis released at https://github.com/ufal/multilexnorm2021 and the fine-tuned\nmodels at https://huggingface.co/ufal.\n

Paper

References (39)

Scroll for more · 27 remaining

Similar papers

© 2026 NYSGPT2525 LLC