ScrapLLM: Benchmarking Data Acquisition for Research in Hate Speech Detection

This study explores the role of Large Language Models (LLMs) in the domain of hate speech detection through dataset generation and analysis. By combining social media scraping with prompt engineering, we developed a benchmark dataset of 1,000 entries encompassing multiple categories of hate speech using ChatGPT as a case study. The methodology demonstrates how LLMs can be harnessed not only for detection but also for synthetic data generation, thereby enriching resources for research. At the same time, the work highlights the risks of adversarial misuse, particularly through jailbreaking prompts that bypass safety filters and elicit harmful outputs. Addressing such vulnerabilities remains a critical challenge for the responsible deployment of LLMs. Ethical concerns-including privacy, consent, and potential misuse-were integrated throughout the study to ensure that the dataset serves constructive purposes only. Overall, this research contributes to responsible AI by providing both technical insights and ethical considerations for LLM-based dataset generation.

Paper

Full text

PDF

ScrapLLM: Benchmarking Data Acquisition for Research in Hate Speech Detection

Semantic Scholar · 2025

Abstract

This study explores the role of Large Language Models (LLMs) in the domain of hate speech detection through dataset generation and analysis. By combining social media scraping with prompt engineering, we developed a benchmark dataset of 1,000 entries encompassing multiple categories of hate speech using ChatGPT as a case study. The methodology demonstrates how LLMs can be harnessed not only for detection but also for synthetic data generation, thereby enriching resources for research. At the same time, the work highlights the risks of adversarial misuse, particularly through jailbreaking prompts that bypass safety filters and elicit harmful outputs. Addressing such vulnerabilities remains a critical challenge for the responsible deployment of LLMs. Ethical concerns-including privacy, consent, and potential misuse-were integrated throughout the study to ensure that the dataset serves constructive purposes only. Overall, this research contributes to responsible AI by providing both technical insights and ethical considerations for LLM-based dataset generation.

Similar papers

© 2026 NYSGPT2525 LLC