Can Crowdsourcing Survive the LLM Era? A Community Survey on Human Data Collection

The widespread use of Large Language Models (LLMs) as writing tools challenges the validity of crowdsourced data, as crowdworkers may outsource tasks to models. To better understand how this is addressed, we surveyed 155 researchers in NLP and related disciplines about their experiences and opinions on collecting free-text responses via crowdsourcing. This paper provides an overview of practitioners'challenges, mitigation strategies, and the foreseen implications on data quality. 44% of respondents reported observing LLM usage in their crowdsourced data. While 93% of them had anticipated this, half were unsure what precautions to take. The most prevalent detection strategies are distinctive textual style patterns and unusually fast completion times. Overall, survey responses show that the research community is aware of the problem and taking measures, but existing efforts remain insufficient to fully address it. Finally, we derive a set of considerations to guide future crowdsourced free-text data collection in the era of LLMs.

Paper

References (16)

102023. Survey of hallucination in natural language generationACM Computing Surveys
11re-shape , scenario-writing to explore diverse implications of generative AI in the news environmentAI and Ethics
12A.3 RQ3 additional material: Examples for acceptable LLM use Table 3 lists the themes identified from responses to question O5, grouped by acceptability category, with representative

Scroll for more · 4 remaining

Similar papers

© 2026 NYSGPT2525 LLC