We present the winning entry to the Multilingual Lexical Normalization\n(MultiLexNorm) shared task at W-NUT 2021 (van der Goot et al., 2021a), which\nevaluates lexical-normalization systems on 12 social media datasets in 11\nlanguages. We base our solution on a pre-trained byte-level language model,\nByT5 (Xue et al., 2021a), which we further pre-train on synthetic data and then\nfine-tune on authentic normalization data. Our system achieves the best\nperformance by a wide margin in intrinsic evaluation, and also the best\nperformance in extrinsic evaluation through dependency parsing. The source code\nis released at https://github.com/ufal/multilexnorm2021 and the fine-tuned\nmodels at https://huggingface.co/ufal.\n
Paper
References (39)
Scroll for more · 27 remaining