DeepHateExplainer: Explainable Hate Speech Detection in Under-resourced Bengali Language

The exponential growths of social media and micro-blogging sites not only\nprovide platforms for empowering freedom of expressions and individual voices,\nbut also enables people to express anti-social behaviour like online\nharassment, cyberbullying, and hate speech. Numerous works have been proposed\nto utilize textual data for social and anti-social behaviour analysis, by\npredicting the contexts mostly for highly-resourced languages like English.\nHowever, some languages are under-resourced, e.g., South Asian languages like\nBengali, that lack computational resources for accurate natural language\nprocessing (NLP). In this paper, we propose an explainable approach for hate\nspeech detection from the under-resourced Bengali language, which we called\nDeepHateExplainer. Bengali texts are first comprehensively preprocessed, before\nclassifying them into political, personal, geopolitical, and religious hates\nusing a neural ensemble method of transformer-based neural architectures (i.e.,\nmonolingual Bangla BERT-base, multilingual BERT-cased/uncased, and\nXLM-RoBERTa). Important(most and least) terms are then identified using\nsensitivity analysis and layer-wise relevance propagation(LRP), before\nproviding human-interpretable explanations. Finally, we compute\ncomprehensiveness and sufficiency scores to measure the quality of explanations\nw.r.t faithfulness. Evaluations against machine learning~(linear and tree-based\nmodels) and neural networks (i.e., CNN, Bi-LSTM, and Conv-LSTM with word\nembeddings) baselines yield F1-scores of 78%, 91%, 89%, and 84%, for political,\npersonal, geopolitical, and religious hates, respectively, outperforming both\nML and DNN baselines.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC