With the continuous advancement of the text generation capabilities of large language models, machine-generated text has found widespread application across various industries. However, this proliferation has also introduced risks associated with the misuse of such text. Consequently, effectively determining the origin of such text has become a significant research topic. However, existing AI text detectors often struggle to achieve high accuracy rates in practical scenarios involving machine-generated text from specific domains and diverse models. This paper uses the Machine Generated Text Detection in the Wild (MAGE) detection platform and its pre-trained language model detector. It conducts domain-specific fine-tuning and experimental validation across three specialised domains (Comment, Q&A and News texts) and a mixed Multi-domains setting. The experimental results demonstrate that, following fine-tuning using domain-specific custom datasets, the detector achieves significantly improved accuracy and other metrics across all three domains and the mixed domain.
Paper
Full text
Research on Large Language Model Text Source Detection Based on Domain Adaptive Fine-Tuning
Semantic Scholar · 2025
Abstract
With the continuous advancement of the text generation capabilities of large language models, machine-generated text has found widespread application across various industries. However, this proliferation has also introduced risks associated with the misuse of such text. Consequently, effectively determining the origin of such text has become a significant research topic. However, existing AI text detectors often struggle to achieve high accuracy rates in practical scenarios involving machine-generated text from specific domains and diverse models. This paper uses the Machine Generated Text Detection in the Wild (MAGE) detection platform and its pre-trained language model detector. It conducts domain-specific fine-tuning and experimental validation across three specialised domains (Comment, Q&A and News texts) and a mixed Multi-domains setting. The experimental results demonstrate that, following fine-tuning using domain-specific custom datasets, the detector achieves significantly improved accuracy and other metrics across all three domains and the mixed domain.