A Systematic Analysis of Oversampling and Classic Machine Learning Pipelines for Multilingual Offensive Language Detection in Resource-Constrained Scenarios

While recent advances in Natural Language Processing (NLP) are increasingly dominated by large-scale deep learning (DL), the detection of offensive language in real-world applications still relies heavily on classical machine learning (ML) approaches due to computational constraints and the need for interpretability. Despite the critical importance of identifying harmful content, existing studies typically focus either on single-language settings or on computationally intensive models, leaving a significant gap in systematic, cross-lingual evaluations of lightweight detection pipelines under resource-constrained conditions. This study addresses this gap by presenting a large-scale benchmark and evaluation of classical ML-based classification pipelines across 15 datasets covering 6 languages. Our findings yield two major insights. First, widely adopted preprocessing techniques often provide minimal performance gains, and in several cases may even degrade detection accuracy. Second, oversampling emerges as the most consistently effective strategy for boosting performance, improving F1-scores significantly in the highly imbalanced datasets typical of toxic content moderation. Overall, this work challenges conventional assumptions in NLP pipeline design and provides practical, evidence-based recommendations for building efficient and robust multilingual offensive language detection systems under real-world constraints.

Paper

Full text

PDF

A Systematic Analysis of Oversampling and Classic Machine Learning Pipelines for Multilingual Offensive Language Detection in Resource-Constrained Scenarios

Semantic Scholar · Computer Science · 2026

Abstract

While recent advances in Natural Language Processing (NLP) are increasingly dominated by large-scale deep learning (DL), the detection of offensive language in real-world applications still relies heavily on classical machine learning (ML) approaches due to computational constraints and the need for interpretability. Despite the critical importance of identifying harmful content, existing studies typically focus either on single-language settings or on computationally intensive models, leaving a significant gap in systematic, cross-lingual evaluations of lightweight detection pipelines under resource-constrained conditions. This study addresses this gap by presenting a large-scale benchmark and evaluation of classical ML-based classification pipelines across 15 datasets covering 6 languages. Our findings yield two major insights. First, widely adopted preprocessing techniques often provide minimal performance gains, and in several cases may even degrade detection accuracy. Second, oversampling emerges as the most consistently effective strategy for boosting performance, improving F1-scores significantly in the highly imbalanced datasets typical of toxic content moderation. Overall, this work challenges conventional assumptions in NLP pipeline design and provides practical, evidence-based recommendations for building efficient and robust multilingual offensive language detection systems under real-world constraints.

References (47)

Scroll for more · 35 remaining

Similar papers

© 2026 NYSGPT2525 LLC