This study explores the role of Large Language Models (LLMs) in the domain of hate speech detection through dataset generation and analysis. By combining social media scraping with prompt engineering, we developed a benchmark dataset of 1,000 entries encompassing multiple categories of hate speech using ChatGPT as a case study. The methodology demonstrates how LLMs can be harnessed not only for detection but also for synthetic data generation, thereby enriching resources for research. At the same time, the work highlights the risks of adversarial misuse, particularly through jailbreaking prompts that bypass safety filters and elicit harmful outputs. Addressing such vulnerabilities remains a critical challenge for the responsible deployment of LLMs. Ethical concerns-including privacy, consent, and potential misuse-were integrated throughout the study to ensure that the dataset serves constructive purposes only. Overall, this research contributes to responsible AI by providing both technical insights and ethical considerations for LLM-based dataset generation.
Paper
Full text
ScrapLLM: Benchmarking Data Acquisition for Research in Hate Speech Detection
Semantic Scholar · 2025
Abstract
This study explores the role of Large Language Models (LLMs) in the domain of hate speech detection through dataset generation and analysis. By combining social media scraping with prompt engineering, we developed a benchmark dataset of 1,000 entries encompassing multiple categories of hate speech using ChatGPT as a case study. The methodology demonstrates how LLMs can be harnessed not only for detection but also for synthetic data generation, thereby enriching resources for research. At the same time, the work highlights the risks of adversarial misuse, particularly through jailbreaking prompts that bypass safety filters and elicit harmful outputs. Addressing such vulnerabilities remains a critical challenge for the responsible deployment of LLMs. Ethical concerns-including privacy, consent, and potential misuse-were integrated throughout the study to ensure that the dataset serves constructive purposes only. Overall, this research contributes to responsible AI by providing both technical insights and ethical considerations for LLM-based dataset generation.