Fine-Grained Detection of Implicit Hate Speech in Chinese Based on Contrastive Learning and Retrieval-Augmented Adjudication

Implicit hate speech is difficult to detect because hostile intent is often conveyed through metaphor, irony, coded expressions, stereotypes, or culturally situated allusions rather than direct insults. We first construct a fine-grained Chinese implicit hate speech dataset containing 19,939 samples from Zhihu and Baidu Tieba, covering nine target categories and three expression labels. Based on this dataset, we propose FICHS, a fine-grained and implicit Chinese hate speech detection framework with three synergistic modules built on a RoBERTa-based base detector: an Explicit–Implicit Supervised Contrastive Module that learns a discriminative semantic space between explicit and implicit hate speech, a Semantic Opacity-Aware Routing and Adjudication module that selectively routes low-confidence but high-implicitness samples to an LLM for secondary judgment, and a Pattern-Guided Retrieval-Augmented Recovery Module that retrieves socio-cultural risk patterns to support second-round review and false-negative recovery. Experimental results demonstrate that FICHS effectively improves Chinese implicit hate speech detection, achieving a weighted F1 score of 0.8391, with precision of 0.8428 and recall of 0.8396, outperforming both the SOTA detector and standalone LLM baselines.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC