Abusive and Threatening Language Detection in Urdu using Boosting based and BERT based models: A Comparative Approach

Online hatred is a growing concern on many social media platforms. To address\nthis issue, different social media platforms have introduced moderation\npolicies for such content. They also employ moderators who can check the posts\nviolating moderation policies and take appropriate action. Academicians in the\nabusive language research domain also perform various studies to detect such\ncontent better. Although there is extensive research in abusive language\ndetection in English, there is a lacuna in abusive language detection in low\nresource languages like Hindi, Urdu etc. In this FIRE 2021 shared task -\n"HASOC- Abusive and Threatening language detection in Urdu" the organizers\npropose an abusive language detection dataset in Urdu along with threatening\nlanguage detection. In this paper, we explored several machine learning models\nsuch as XGboost, LGBM, m-BERT based models for abusive and threatening content\ndetection in Urdu based on the shared task. We observed the Transformer model\nspecifically trained on abusive language dataset in Arabic helps in getting the\nbest performance. Our model came First for both abusive and threatening content\ndetection with an F1scoreof 0.88 and 0.54, respectively.\n

Paper

References (26)

Scroll for more · 14 remaining

Similar papers

© 2026 NYSGPT2525 LLC