GUIR at SemEval-2020 Task 12: Domain-Tuned Contextualized Models for Offensive Language Detection

Offensive language detection is an important and challenging task in natural\nlanguage processing. We present our submissions to the OffensEval 2020 shared\ntask, which includes three English sub-tasks: identifying the presence of\noffensive language (Sub-task A), identifying the presence of target in\noffensive language (Sub-task B), and identifying the categories of the target\n(Sub-task C). Our experiments explore using a domain-tuned contextualized\nlanguage model (namely, BERT) for this task. We also experiment with different\ncomponents and configurations (e.g., a multi-view SVM) stacked upon BERT models\nfor specific sub-tasks. Our submissions achieve F1 scores of 91.7% in Sub-task\nA, 66.5% in Sub-task B, and 63.2% in Sub-task C. We perform an ablation study\nwhich reveals that domain tuning considerably improves the classification\nperformance. Furthermore, error analysis shows common misclassification errors\nmade by our model and outlines research directions for future.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC