Social media platforms have grown so rapidly that the volume of content generated by users has skyrocketed, and at the same time, the toxic, hateful, offensive, and abusive languages have become very common online. The problem has been getting worse to the point that researchers have felt the need to design advanced computing systems that can, with high efficiency and dependability, detect toxic comments automatically. The review examines yearly research articles that use transformer-based language models, hybrid architectural systems, machine learning (ML) and deep learning (DL) techniques to identify harmful comments in a variety of languages. Researchers developed sophisticated neural network models as a result of earlier studies that used conventional algorithms like Random Forests, Support Vector Machines (SVM), and Logistic Regression. These studies demonstrated that the systems lacked sufficient context comprehension. Neural networks have become more powerful, including CNNs, LSTMs, Bidirectional LSTMs, Gated Recurrent Units, and the attention mechanism, and thus their use in recent studies has become widespread. The research team aims to create a system which combines the best features of transformer-based language frameworks with their pre-trained BERT and RoBERTA and DistilBERT models because these advanced systems outperform traditional methods through their ability to understand context better and utilize transfer learning. The multilingual and low-resource language studies encounter obstacles which stem from code-mixing and spelling variations and the lack of extensive annotated datasets. The research fields of multi-label toxicity classification and adversarial robustness and interpretability frameworks currently receive research interest while researchers work to solve the problems that emerge during actual system operation. The review examines two main research directions which focus on developing toxic comment detection systems for contemporary social media platforms which require large-scale development and ethical practices and multilingual capability.
Paper
Full text
Threatshield: Toxic Comments on Social Media
Semantic Scholar · 2026
Abstract
Social media platforms have grown so rapidly that the volume of content generated by users has skyrocketed, and at the same time, the toxic, hateful, offensive, and abusive languages have become very common online. The problem has been getting worse to the point that researchers have felt the need to design advanced computing systems that can, with high efficiency and dependability, detect toxic comments automatically. The review examines yearly research articles that use transformer-based language models, hybrid architectural systems, machine learning (ML) and deep learning (DL) techniques to identify harmful comments in a variety of languages. Researchers developed sophisticated neural network models as a result of earlier studies that used conventional algorithms like Random Forests, Support Vector Machines (SVM), and Logistic Regression. These studies demonstrated that the systems lacked sufficient context comprehension. Neural networks have become more powerful, including CNNs, LSTMs, Bidirectional LSTMs, Gated Recurrent Units, and the attention mechanism, and thus their use in recent studies has become widespread. The research team aims to create a system which combines the best features of transformer-based language frameworks with their pre-trained BERT and RoBERTA and DistilBERT models because these advanced systems outperform traditional methods through their ability to understand context better and utilize transfer learning. The multilingual and low-resource language studies encounter obstacles which stem from code-mixing and spelling variations and the lack of extensive annotated datasets. The research fields of multi-label toxicity classification and adversarial robustness and interpretability frameworks currently receive research interest while researchers work to solve the problems that emerge during actual system operation. The review examines two main research directions which focus on developing toxic comment detection systems for contemporary social media platforms which require large-scale development and ethical practices and multilingual capability.