Exploring Transformer Based Models to Identify Hate Speech and Offensive Content in English and Indo-Aryan Languages
Hate speech is considered to be one of the major issues currently plaguing\nonline social media. Repeated and repetitive exposure to hate speech has been\nshown to create physiological effects on the target users. Thus, hate speech,\nin all its forms, should be addressed on these platforms in order to maintain\ngood health. In this paper, we explored several Transformer based machine\nlearning models for the detection of hate speech and offensive content in\nEnglish and Indo-Aryan languages at FIRE 2021. We explore several models such\nas mBERT, XLMR-large, XLMR-base by team name "Super Mario". Our models came 2nd\nposition in Code-Mixed Data set (Macro F1: 0.7107), 2nd position in Hindi\ntwo-class classification(Macro F1: 0.7797), 4th in English four-class category\n(Macro F1: 0.8006) and 12th in English two-class category (Macro F1: 0.6447).\n