Research on Chinese multimodal fake review detection for AIGC-Generated content

To address the lack of Chinese benchmark resources for detecting AI-generated (AIGC) multimodal fake reviews, we construct and validate a large-scale text–image dataset and benchmark to support platform governance and consumer protection. Authentic restaurant reviews were curated and paired with synthetic deceptive counterparts generated by a large language–vision model, yielding a balanced Chinese multimodal dataset of over 20,000 text–image samples. Mainstream unimodal (text-only, image-only) and multimodal pre-trained models were evaluated under a unified protocol. We further conducted generalization tests via information-perturbation stress tests and cross-lingual transfer scenarios, and compared fusion strategies (early, intermediate, deep). Multimodal models employing deep fusion consistently outperform unimodal and shallow-fusion baselines in accuracy and robustness. They retain superior performance under feature perturbations and demonstrate stronger transferability across languages, confirming the benefit of jointly leveraging complementary textual and visual cues for fake-review detection. This work presents, to our knowledge, the first Chinese multimodal AIGC fake review dataset accompanied by a comprehensive benchmark. It provides an open, reproducible resource and empirical evidence that deep multimodal fusion substantially improves detection effectiveness and robustness, offering practical guidance for future research and real-world deployment.

Paper

Full text

PDF

Research on Chinese multimodal fake review detection for AIGC-Generated content

Semantic Scholar · 2026

Abstract

To address the lack of Chinese benchmark resources for detecting AI-generated (AIGC) multimodal fake reviews, we construct and validate a large-scale text–image dataset and benchmark to support platform governance and consumer protection.

Authentic restaurant reviews were curated and paired with synthetic deceptive counterparts generated by a large language–vision model, yielding a balanced Chinese multimodal dataset of over 20,000 text–image samples. Mainstream unimodal (text-only, image-only) and multimodal pre-trained models were evaluated under a unified protocol. We further conducted generalization tests via information-perturbation stress tests and cross-lingual transfer scenarios, and compared fusion strategies (early, intermediate, deep).

Multimodal models employing deep fusion consistently outperform unimodal and shallow-fusion baselines in accuracy and robustness. They retain superior performance under feature perturbations and demonstrate stronger transferability across languages, confirming the benefit of jointly leveraging complementary textual and visual cues for fake-review detection.

This work presents, to our knowledge, the first Chinese multimodal AIGC fake review dataset accompanied by a comprehensive benchmark. It provides an open, reproducible resource and empirical evidence that deep multimodal fusion substantially improves detection effectiveness and robustness, offering practical guidance for future research and real-world deployment.

Similar papers

© 2026 NYSGPT2525 LLC