An Information Retrieval Approach to Building Datasets for Hate Speech Detection

Building a benchmark dataset for hate speech detection presents various\nchallenges. Firstly, because hate speech is relatively rare, random sampling of\ntweets to annotate is very inefficient in finding hate speech. To address this,\nprior datasets often include only tweets matching known "hate words". However,\nrestricting data to a pre-defined vocabulary may exclude portions of the\nreal-world phenomenon we seek to model. A second challenge is that definitions\nof hate speech tend to be highly varying and subjective. Annotators having\ndiverse prior notions of hate speech may not only disagree with one another but\nalso struggle to conform to specified labeling guidelines. Our key insight is\nthat the rarity and subjectivity of hate speech are akin to that of relevance\nin information retrieval (IR). This connection suggests that well-established\nmethodologies for creating IR test collections can be usefully applied to\ncreate better benchmark datasets for hate speech. To intelligently and\nefficiently select which tweets to annotate, we apply standard IR techniques of\n{\\em pooling} and {\\em active learning}. To improve both consistency and value\nof annotations, we apply {\\em task decomposition} and {\\em annotator rationale}\ntechniques. We share a new benchmark dataset for hate speech detection on\nTwitter that provides broader coverage of hate than prior datasets. We also\nshow a dramatic drop in accuracy of existing detection models when tested on\nthese broader forms of hate. Annotator rationales we collect not only justify\nlabeling decisions but also enable future work opportunities for\ndual-supervision and/or explanation generation in modeling. Further details of\nour approach can be found in the supplementary materials.\n

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC