Leveraging cross-platform data to improve automated hate speech detection

Hate speech is increasingly prevalent online, and its negative outcomes\ninclude increased prejudice, extremism, and even offline hate crime. Automatic\ndetection of online hate speech can help us to better understand these impacts.\nHowever, while the field has recently progressed through advances in natural\nlanguage processing, challenges still remain. In particular, most existing\napproaches for hate speech detection focus on a single social media platform in\nisolation. This limits both the use of these models and their validity, as the\nnature of language varies from platform to platform. Here we propose a new\ncross-platform approach to detect hate speech which leverages multiple datasets\nand classification models from different platforms and trains a superlearner\nthat can combine existing and novel training data to improve detection and\nincrease model applicability. We demonstrate how this approach outperforms\nexisting models, and achieves good performance when tested on messages from\nnovel social media platforms not included in the original training data.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC