A Systematic Analysis of Oversampling and Classic Machine Learning Pipelines for Multilingual Offensive Language Detection in Resource-Constrained Scenarios
While recent advances in Natural Language Processing (NLP) are increasingly dominated by large-scale deep learning (DL), the detection of offensive language in real-world applications still relies heavily on classical machine learning (ML) approaches due to computational constraints and the need for interpretability. Despite the critical importance of identifying harmful content, existing studies typically focus either on single-language settings or on computationally intensive models, leaving a significant gap in systematic, cross-lingual evaluations of lightweight detection pipelines under resource-constrained conditions. This study addresses this gap by presenting a large-scale benchmark and evaluation of classical ML-based classification pipelines across 15 datasets covering 6 languages. Our findings yield two major insights. First, widely adopted preprocessing techniques often provide minimal performance gains, and in several cases may even degrade detection accuracy. Second, oversampling emerges as the most consistently effective strategy for boosting performance, improving F1-scores significantly in the highly imbalanced datasets typical of toxic content moderation. Overall, this work challenges conventional assumptions in NLP pipeline design and provides practical, evidence-based recommendations for building efficient and robust multilingual offensive language detection systems under real-world constraints.
Paper
Full text
A Systematic Analysis of Oversampling and Classic Machine Learning Pipelines for Multilingual Offensive Language Detection in Resource-Constrained Scenarios
Semantic Scholar · Computer Science · 2026
Abstract
While recent advances in Natural Language Processing (NLP) are increasingly dominated by large-scale deep learning (DL), the detection of offensive language in real-world applications still relies heavily on classical machine learning (ML) approaches due to computational constraints and the need for interpretability. Despite the critical importance of identifying harmful content, existing studies typically focus either on single-language settings or on computationally intensive models, leaving a significant gap in systematic, cross-lingual evaluations of lightweight detection pipelines under resource-constrained conditions. This study addresses this gap by presenting a large-scale benchmark and evaluation of classical ML-based classification pipelines across 15 datasets covering 6 languages. Our findings yield two major insights. First, widely adopted preprocessing techniques often provide minimal performance gains, and in several cases may even degrade detection accuracy. Second, oversampling emerges as the most consistently effective strategy for boosting performance, improving F1-scores significantly in the highly imbalanced datasets typical of toxic content moderation. Overall, this work challenges conventional assumptions in NLP pipeline design and provides practical, evidence-based recommendations for building efficient and robust multilingual offensive language detection systems under real-world constraints.
References (47)
Scroll for more · 35 remaining