Fine-tuning of Pre-trained Transformers for Hate, Offensive, and Profane Content Detection in English and Marathi
This paper describes neural models developed for the Hate Speech and\nOffensive Content Identification in English and Indo-Aryan Languages Shared\nTask 2021. Our team called neuro-utmn-thales participated in two tasks on\nbinary and fine-grained classification of English tweets that contain hate,\noffensive, and profane content (English Subtasks A & B) and one task on\nidentification of problematic content in Marathi (Marathi Subtask A). For\nEnglish subtasks, we investigate the impact of additional corpora for hate\nspeech detection to fine-tune transformer models. We also apply a one-vs-rest\napproach based on Twitter-RoBERTa to discrimination between hate, profane and\noffensive posts. Our models ranked third in English Subtask A with the F1-score\nof 81.99% and ranked second in English Subtask B with the F1-score of 65.77%.\nFor the Marathi tasks, we propose a system based on the Language-Agnostic BERT\nSentence Embedding (LaBSE). This model achieved the second result in Marathi\nSubtask A obtaining an F1 of 88.08%.\n