Leveraging Multi-domain, Heterogeneous Data using Deep Multitask Learning for Hate Speech Detection
With the exponential rise in user-generated web content on social media, the\nproliferation of abusive languages towards an individual or a group across the\ndifferent sections of the internet is also rapidly increasing. It is very\nchallenging for human moderators to identify the offensive contents and filter\nthose out. Deep neural networks have shown promise with reasonable accuracy for\nhate speech detection and allied applications. However, the classifiers are\nheavily dependent on the size and quality of the training data. Such a\nhigh-quality large data set is not easy to obtain. Moreover, the existing data\nsets that have emerged in recent times are not created following the same\nannotation guidelines and are often concerned with different types and\nsub-types related to hate. To solve this data sparsity problem, and to obtain\nmore global representative features, we propose a Convolution Neural Network\n(CNN) based multi-task learning models (MTLs)\\footnote{code is available at\nhttps://github.com/imprasshant/STL-MTL} to leverage information from multiple\nsources. Empirical analysis performed on three benchmark datasets shows the\nefficacy of the proposed approach with the significant improvement in accuracy\nand F-score to obtain state-of-the-art performance with respect to the existing\nsystems.\n