KBCNMUJAL@HASOC-Dravidian-CodeMix-FIRE2020: Using Machine Learning for Detection of Hate Speech and Offensive Code-Mixed Social Media text
This paper describes the system submitted by our team, KBCNMUJAL, for Task 2\nof the shared task Hate Speech and Offensive Content Identification in\nIndo-European Languages (HASOC), at Forum for Information Retrieval Evaluation,\nDecember 16-20, 2020, Hyderabad, India. The datasets of two Dravidian languages\nViz. Malayalam and Tamil of size 4000 observations, each were shared by the\nHASOC organizers. These datasets are used to train the machine using different\nmachine learning algorithms, based on classification and regression models. The\ndatasets consist of tweets or YouTube comments with two class labels offensive\nand not offensive. The machine is trained to classify such social media\nmessages in these two categories. Appropriate n-gram feature sets are extracted\nto learn the specific characteristics of the Hate Speech text messages. These\nfeature models are based on TFIDF weights of n-gram. The referred work and\nrespective experiments show that the features such as word, character and\ncombined model of word and character n-grams could be used to identify the term\npatterns of offensive text contents. As a part of the HASOC shared task, the\ntest data sets are made available by the HASOC track organizers. The best\nperforming classification models developed for both languages are applied on\ntest datasets. The model which gives the highest accuracy result on training\ndataset for Malayalam language was experimented to predict the categories of\nrespective test data. This system has obtained an F1 score of 0.77. Similarly\nthe best performing model for Tamil language has obtained an F1 score of 0.87.\nThis work has received 2nd and 3rd rank in this shared Task 2 for Malayalam and\nTamil language respectively. The proposed system is named HASOC_kbcnmujal.\n
Paper
References (17)
Scroll for more · 5 remaining