ExBERT-CM: Explainable Cyberbullying Detection in Code-Mixed Social Media with Multilingual Transformer Models
On social media, cyberbullying is particularly an alarming effect on mental health, as in multilingual cultures the vast use of code-mixed communication is a common practice. Current detection systems are generally monolingual black boxes, which restrict their own applicability and interpretability to mixed-language situations. In this paper, a suggestible cyberbullying detection system to code-mixed social media text is described, as a mixture of state-of-the-art transformers multilingual and post-hoc interpretability systems. Code-mixed datasets in Hindi to English and Tamil to English and other systems publicly packed, are strictly preprocessed and normalised to allow the treatment of transliteration, spelling variation and tokenisation. Multi-language transformer architectures (mBERT, XLM-R and MuRIL) are fine-tuned and ensembled to do binary and multi-class cyberbullying classification. To provide even more transparency, SHAP and attention-based visualisation are included, showing the linguistic basis of each prediction, which improves moderator trust and helps to use AI responsibly. Experimental findings demonstrate that the proposed model is much more effective when applied to the traditional machine-learning baselines, or one specific transformer model since it achieves the overall accuracy of 94.5 % when applied to the benchmark datasets of code-mixed text. Qualitative analysis also proves that the obtained explanations embrace the abusive patterns of mixed-language posts facilitating more effective and less biased content management. This paper provides one of the first high-fidelity and interpretable cyberbullying detection architectures in the lack of study under-representative multi-language environments and frameworks explainability as one significant production of credible AI application to online social networks.
Paper
Full text
ExBERT-CM: Explainable Cyberbullying Detection in Code-Mixed Social Media with Multilingual Transformer Models
Semantic Scholar · 2025
Abstract
On social media, cyberbullying is particularly an alarming effect on mental health, as in multilingual cultures the vast use of code-mixed communication is a common practice. Current detection systems are generally monolingual black boxes, which restrict their own applicability and interpretability to mixed-language situations. In this paper, a suggestible cyberbullying detection system to code-mixed social media text is described, as a mixture of state-of-the-art transformers multilingual and post-hoc interpretability systems. Code-mixed datasets in Hindi to English and Tamil to English and other systems publicly packed, are strictly preprocessed and normalised to allow the treatment of transliteration, spelling variation and tokenisation. Multi-language transformer architectures (mBERT, XLM-R and MuRIL) are fine-tuned and ensembled to do binary and multi-class cyberbullying classification. To provide even more transparency, SHAP and attention-based visualisation are included, showing the linguistic basis of each prediction, which improves moderator trust and helps to use AI responsibly. Experimental findings demonstrate that the proposed model is much more effective when applied to the traditional machine-learning baselines, or one specific transformer model since it achieves the overall accuracy of 94.5 % when applied to the benchmark datasets of code-mixed text. Qualitative analysis also proves that the obtained explanations embrace the abusive patterns of mixed-language posts facilitating more effective and less biased content management. This paper provides one of the first high-fidelity and interpretable cyberbullying detection architectures in the lack of study under-representative multi-language environments and frameworks explainability as one significant production of credible AI application to online social networks.