Countering hate on social media: Large scale classification of hate and counter speech

Hateful rhetoric is plaguing online discourse, fostering extreme societal\nmovements and possibly giving rise to real-world violence. A potential solution\nto this growing global problem is citizen-generated counter speech where\ncitizens actively engage in hate-filled conversations to attempt to restore\ncivil non-polarized discourse. However, its actual effectiveness in curbing the\nspread of hatred is unknown and hard to quantify. One major obstacle to\nresearching this question is a lack of large labeled data sets for training\nautomated classifiers to identify counter speech. Here we made use of a unique\nsituation in Germany where self-labeling groups engaged in organized online\nhate and counter speech. We used an ensemble learning algorithm which pairs a\nvariety of paragraph embeddings with regularized logistic regression functions\nto classify both hate and counter speech in a corpus of millions of relevant\ntweets from these two groups. Our pipeline achieved macro F1 scores on out of\nsample balanced test sets ranging from 0.76 to 0.97---accuracy in line and even\nexceeding the state of the art. On thousands of tweets, we used crowdsourcing\nto verify that the judgments made by the classifier are in close alignment with\nhuman judgment. We then used the classifier to discover hate and counter speech\nin more than 135,000 fully-resolved Twitter conversations occurring from 2013\nto 2018 and study their frequency and interaction. Altogether, our results\nhighlight the potential of automated methods to evaluate the impact of\ncoordinated counter speech in stabilizing conversations on social media.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC