AI-UPV at IberLEF-2021 DETOXIS task: Toxicity Detection in Immigration-Related Web News Comments Using Transformers and Statistical Models

This paper describes our participation in the DEtection of TOXicity in\ncomments In Spanish (DETOXIS) shared task 2021 at the 3rd Workshop on Iberian\nLanguages Evaluation Forum. The shared task is divided into two related\nclassification tasks: (i) Task 1: toxicity detection and; (ii) Task 2: toxicity\nlevel detection. They focus on the xenophobic problem exacerbated by the spread\nof toxic comments posted in different online news articles related to\nimmigration. One of the necessary efforts towards mitigating this problem is to\ndetect toxicity in the comments. Our main objective was to implement an\naccurate model to detect xenophobia in comments about web news articles within\nthe DETOXIS shared task 2021, based on the competition's official metrics: the\nF1-score for Task 1 and the Closeness Evaluation Metric (CEM) for Task 2. To\nsolve the tasks, we worked with two types of machine learning models: (i)\nstatistical models and (ii) Deep Bidirectional Transformers for Language\nUnderstanding (BERT) models. We obtained our best results in both tasks using\nBETO, an BERT model trained on a big Spanish corpus. We obtained the 3rd place\nin Task 1 official ranking with the F1-score of 0.5996, and we achieved the 6th\nplace in Task 2 official ranking with the CEM of 0.7142. Our results suggest:\n(i) BERT models obtain better results than statistical models for toxicity\ndetection in text comments; (ii) Monolingual BERT models have an advantage over\nmultilingual BERT models in toxicity detection in text comments in their\npre-trained language.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC