Summary
The paper presents SAID, a benchmark to assess the AI-generated text that is shared on two online platforms: Quora and Zhihu. The paper releases a large-scale dataset of English and Chinese posts that are either AI-generated or human-generated (based on labels or platform features from Quora and Zhihu). Also, the paper investigates whether humans can identify AI-generated text, what is the performance of existing detectors in identifying AI-generated text, as well as designing a classification task that considers both content and user information to classify content as AI-generated content or human-generated content.
Strengths
The paper focuses on an important and timely issue that I believe will become even more important in the future, so these research efforts and benchmarks can equip the community with the necessary tools to being able to identify machine-generated text. An important strength of the paper is that it releases a large-scale dataset of AI and human-generated posts shared on two social media platforms across two languages (English and Chinese). The multilingual and multi-platform aspect of the released dataset is worth noting, given that it likely increases the probability of being used by other researchers focusing on similar topics. Finally, I find the idea of incorporating both content information and user information to identify AI-generated content interesting and it provides a new perspective related to AI-generated content that is particularly shared on social media platforms like Quora and Zhihu.
Weaknesses
I think that the paper has some considerable flaws regarding the methods/assumptions made by the paper, the interpretation of the results, the lack of validation of the employed methods, and the lack of scope. I elaborate on my main concerns along with some suggestions on how to improve the manuscript below.
**Validation and Assumptions:** I think that the paper uses methods for identifying machine-generated text that are not properly validated. Specifically, the paper uses labels provided by Zhihu and the collapsing feature in Quora to determine if the text is machine-generated or not. The issue is that these methods are essentially black boxes, and we do not know how accurate these methods are in determining whether text is machine-generated or not. Also, we do not have any clue how generalizable the methods are and how they can get deceived depending on the prompt that a user will use to generate text (e.g., a user leveraging a prompt that aims to generate text that looks like it is written by a human). On top of this lack of validation of the methods, the paper makes an assumption that any text posted by users that have many posts with these labels are actually generated by machines. Again this is an assumption that is not validated by the paper and it is unclear how the lack of the validation of methods and the assumptions made are likely to affect the quality of the released dataset. A potential way to mitigate these issues is to make a manual validation of the a small sample of the dataset to verify how many posts are indeed machine-generated based on the perceptions of the manual annotators. This help us understand the quality of the dataset and shed light into the potential false positives that exist in the released dataset. This is of paramount importance as the paper aims to release a benchmark dataset that will be used by other researchers, hence the quality of the dataset is very important.
**Sample sizes and interpretation of the results:** Another important concern with the paper is that the results presented in Section 4 are based on the perspectives of two participants only, hence the results do not provide much confidence. Also, the interpretation of the results and the comparisons with the OpenAI work is not fair given that the previous work used non-expert users, whereas here, you use participants who are familiar with LLMs and social media. Overall, I suggest to the authors to provide more details on the recruited participants (e.g., are participants authors of this paper?) and extend the annotation process so that it includes more annotators. I believe that making generalizable conclusions based on two participants is not ideal to say the least. Also, I suggest rephrasing the interpretation of these results and comparisons with previous work to reflect that here you are using expert participants while previous work did not use expert participants (which might likely be one of the reasons for the observed differences in the accuracy). Also, I am a little bit puzzled on why the paper elected to use samples and not the entire dataset for the analysis in Section 6. Using the entire dataset will give a more comprehensive view of the dataset and provide more holistic results. I suggest either using the entire dataset or justifying the need to sample the dataset.
**Scope and Focus of the Paper:** Overall, I believe that the paper tries to study too many aspects of the problem, which results in undertaking a very shallow and not validated analysis of multiple aspects of the problem. I suggest to the authors to limit the scope of the paper and focus on some of the aspects and look at them deeper and ensure that the employed methods are validated and can provide confident results. For instance, I believe that as the paper is, Section 4 is very shallow and does not provide confident results.
Overall, given the above-mentioned concerns, I believe that the paper is not ready for publication yet. I encourage the authors to continue working on this important topic and improve their manuscript and the released dataset.
Questions
1. How accurate are the Zhihu labels and how robust are to machine-generated text that aims to avoid detection (based on the user prompt)?
2. How accurate is the detection of machine-generated text when using the collapse feature in Quora?
3. What is the prevalence of false positives that exist in the released dataset (i..e, text that is considered machine generated when it’s not simply because of the used methods or the user-based assumptions made)?
Rating
3: reject, not good enough
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.