Large Language Models (LLMs) have shown impressive performance in hate speech detection across multiple domains and languages. However, the reliability of their annotations remains an open research question. In particular, their tendency to exhibit self-contradiction producing inconsistent or conflicting classifications for similar or logically related inputs poses a significant challenge to the trustworthiness and interpretability of automated hate speech moderation systems. This paper investigates self-contradictory behaviors in four LLMs: GPT-4o-mini, GPT-4.1-nano, DeepSeek-R1-14B, and Llama-3.1-8B when applied to hate speech detection tasks. We systematically analyze different types of inconsistencies, including content-based contradictions, label contradictions, and contextual drift, across two different datasets containing implicit and explicit hate speech. Our findings reveal that while all models demonstrate high overall accuracy, they frequently contradict themselves in borderline or context-sensitive cases. The study provides critical insights into how LLM variants trade off efficiency with reliability and proposes a framework for identifying and mitigating self-contradiction in hate speech detection pipelines. The LLM responses with anonymized link can be found here: https://github.com/AmitDasRup123/LLM_Self_Contradiction/.
Paper
Full text
Investigating Self-Contradiction in Large Language Models for Hate Speech Detection
Semantic Scholar · 2026
Abstract
Large Language Models (LLMs) have shown impressive performance in hate speech detection across multiple domains and languages. However, the reliability of their annotations remains an open research question. In particular, their tendency to exhibit self-contradiction producing inconsistent or conflicting classifications for similar or logically related inputs poses a significant challenge to the trustworthiness and interpretability of automated hate speech moderation systems. This paper investigates self-contradictory behaviors in four LLMs: GPT-4o-mini, GPT-4.1-nano, DeepSeek-R1-14B, and Llama-3.1-8B when applied to hate speech detection tasks. We systematically analyze different types of inconsistencies, including content-based contradictions, label contradictions, and contextual drift, across two different datasets containing implicit and explicit hate speech. Our findings reveal that while all models demonstrate high overall accuracy, they frequently contradict themselves in borderline or context-sensitive cases. The study provides critical insights into how LLM variants trade off efficiency with reliability and proposes a framework for identifying and mitigating self-contradiction in hate speech detection pipelines. The LLM responses with anonymized link can be found here: https://github.com/AmitDasRup123/LLM\_Self\_Contradiction/.