Cisco at SemEval-2021 Task 5: What's Toxic?: Leveraging Transformers for Multiple Toxic Span Extraction from Online Comments

Social network platforms are generally used to share positive, constructive,\nand insightful content. However, in recent times, people often get exposed to\nobjectionable content like threat, identity attacks, hate speech, insults,\nobscene texts, offensive remarks or bullying. Existing work on toxic speech\ndetection focuses on binary classification or on differentiating toxic speech\namong a small set of categories. This paper describes the system proposed by\nteam Cisco for SemEval-2021 Task 5: Toxic Spans Detection, the first shared\ntask focusing on detecting the spans in the text that attribute to its\ntoxicity, in English language. We approach this problem primarily in two ways:\na sequence tagging approach and a dependency parsing approach. In our sequence\ntagging approach we tag each token in a sentence under a particular tagging\nscheme. Our best performing architecture in this approach also proved to be our\nbest performing architecture overall with an F1 score of 0.6922, thereby\nplacing us 7th on the final evaluation phase leaderboard. We also explore a\ndependency parsing approach where we extract spans from the input sentence\nunder the supervision of target span boundaries and rank our spans using a\nbiaffine model. Finally, we also provide a detailed analysis of our results and\nmodel performance in our paper.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC