Official Comment by Authors
Thank you again for the time and expertise you have invested in these reviews. We appreciate the opportunity to address the concerns and questions raised.
- **(W1) & (W2):** Unless the Ghostbuster benchmark is a subset of the dataset, the Ghostbuster benchmark is considered another dataset and we cannot tell which one is better. The statement "offering a more rigorous evaluation standard." is not verified in the paper. Please clarify why the author(s) consider it.
**Response:** Dear reviewer, we sincerely apologize for comparing the benchmark merely through its domain, which is not full-considered. Since the Ghostbuster benchmark is not a subset of the our benchmark, it is meaningful to evaluate detectors on the Ghostbuster benchmark as well. Therefore, we will add the results on the Ghostbuster benchmark in the final version, and thanks again for your valuable suggestions.
And the statement 'offering a more rigorous evaluation standard' refers to the statement that the benchmark used in our paper includes three large language models. Since cross-model detection is a challenging problem in generated text detection[1,2], we considered our benchmark is more rigorous in this aspect. From a strict perspective, we should not compare benchmark merely depending on the adopted LLMs. Therefore, we will add the results on the Ghostbuster benchmark in the final version, and not compare the benchmarks.
---
- **(W3):** This concern is not addressed. The feature extraction is effective with a certain cost.
**Response:** Dear Reviewer, since we presented comparisons with the methods you mentioned in Weaknesses 1 & 2 of our rebuttal, we did not directly respond to the statement in Weakness 3: 'It is not clear if the feature extraction offers significantly better benefits compared to more lightweight approaches (e.g., [Ref 1][Ref 2] above).' We apologize for any confusion this may have caused and appreciate the opportunity to clarify further.
As shown in Table 1 of the attached PDF, we have compared our results with other methods, including the two lightweight methods you mentioned: Fingerprints [3], and Smaller Models [4]. To facilitate your review, we present the results below. Our method indeed shows a significant improvement over the lightweight methods in terms of average detection AUROC on both our benchmark and Ghostbuster benchmark.
I hope this explanation addresses your concern. If you have any further concerns or questions about our work, we are happy to discuss them with you. We will also add this part to strengthen the manuscript. Thank you again for the time and expertise you have invested in these reviews.
| | ChatGPT | | | | GPT4 | | | | Claude3 | | | | Ghostbuster | | | |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| | XSum | Writing | PubMed | Avg. | XSum | Writing | PubMed | Avg. | XSum | Writing | PubMed | Avg. | News | Creative Writing | Student Essays | Avg. |
| Fingerprints | 0.8815 | 0.8073 | 0.6816 | 0.7901 | 0.8124 | 0.7896 | 0.7133 | 0.7718 | 0.8849 | 0.9033 | 0.8692 | 0.8858 | 0.8108 | 0.8373 | 0.7846 | 0.8109 |
| Smaller Models | 0.9835 | 0.9713 | 0.8843 | 0.9464 | 0.8818 | 0.9098 | 0.8234 | 0.8717 | 0.9798 | 0.9594 | 0.8868 | 0.9420 | **0.9983** | 0.8957 | 0.9673 | 0.9537 |
| DPIC | **1.0000** | **0.9821** | **0.9082** | **0.9634** | **0.9996** | **0.9768** | **0.9438** | **0.9734** | **1.0000** | **0.9950** | **0.9686** | **0.9879** | 0.9950 | **0.9978** | **0.9774** | **0.9900** |
## **References**
[1] Bao G, Zhao Y, Teng Z, et al. Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature[C]//The Twelfth International Conference on Learning Representations.
[2] Yang X, Cheng W, Wu Y, et al. DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text[C]//The Twelfth International Conference on Learning Representations.
[3] McGovern H, Stureborg R, Suhara Y, et al. Your Large Language Models Are Leaving Fingerprints[J]. arXiv preprint arXiv:2405.14057, 2024.
[4] Mireshghallah N, Mattern J, Gao S, et al. Smaller Language Models are Better Zero-shot Machine-Generated Text Detectors[C]//Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers). 2024: 278-293.