The real-world impact of polarization and toxicity in the online sphere\nmarked the end of 2020 and the beginning of this year in a negative way.\nSemeval-2021, Task 5 - Toxic Spans Detection is based on a novel annotation of\na subset of the Jigsaw Unintended Bias dataset and is the first language\ntoxicity detection task dedicated to identifying the toxicity-level spans. For\nthis task, participants had to automatically detect character spans in short\ncomments that render the message as toxic. Our model considers applying Virtual\nAdversarial Training in a semi-supervised setting during the fine-tuning\nprocess of several Transformer-based models (i.e., BERT and RoBERTa), in\ncombination with Conditional Random Fields. Our approach leads to performance\nimprovements and more robust models, enabling us to achieve an F1-score of\n65.73% in the official submission and an F1-score of 66.13% after further\ntuning during post-evaluation.\n
Paper
References (47)
Scroll for more · 35 remaining