Mapping the Increasing Use of LLMs in Scientific Papers

Scientific publishing lays the foundation of science by disseminating research findings, fostering collaboration, encouraging reproducibility, and ensuring that scientific knowledge is accessible, verifiable, and built upon over time. Recently, there has been immense speculation about how many people are using large language models (LLMs) like ChatGPT in their academic writing, and to what extent this tool might have an effect on global scientific practices. However, we lack a precise measure of the proportion of academic writing substantially modified or produced by LLMs. To address this gap, we conduct the first systematic, large-scale analysis across 950,965 papers published between January 2020 and February 2024 on the arXiv, bioRxiv, and Nature portfolio journals, using a population-level statistical framework to measure the prevalence of LLM-modified content over time. Our statistical estimation operates on the corpus level and is more robust than inference on individual instances. Our findings reveal a steady increase in LLM usage, with the largest and fastest growth observed in Computer Science papers (up to 17.5%). In comparison, Mathematics papers and the Nature portfolio showed the least LLM modification (up to 6.3%). Moreover, at an aggregate level, our analysis reveals that higher levels of LLM-modification are associated with papers whose first authors post preprints more frequently, papers in more crowded research areas, and papers of shorter lengths. Our findings suggests that LLMs are being broadly used in scientific writings.

Paper

Similar papers

Reviewer i8UX6/10 · confidence 4/52024-04-30

Summary

Thanks to the authors for the hard work on this paper. The work is well-written and well-argued. It is a relatively small contribution: just applying an existing algorithm to existing data after validating/training with synthetic + real data. The analyses are nice, but can be combined into a single, more comprehensive one (I discuss this in more detail below). In summary, I like this work, but it should be expanded.

Rating

6

Confidence

4

Ethics flag

1

Reasons to accept

The paper is clear. The conclusions make sense and are compelling. The existing analysis seems well-done. Nice charts.

Reasons to reject

- [Minor] Page 5: "This approach also simulates how scientists may be using LLMs..." Need a citation here or another piece of evidence that supports your two-stage approach. - [Minor] Figure 3 is hard to read. I recommend changing y-axis to "error" instead. That's the information I want to extract in any case. - [Major] Sections 5.2, 5.3, and 5.4 are all done independently and in somewhat arbitrary ways. I.e. for 5.2, paper are divided as two or fewer preprints vs three or more, and paper length as below or above 5000. These cutoffs don't seem principally motivated, and the results may be different if they change. Additionally, all of the analyses are done independently. For these kinds of analyses, I recommend instead a fixed-effects model where the output (estimated alpha) is modeled as an output of multiple input factors simultaneously. E.g.: datetime (bucketed), # of preprints (bucketed or not), arxiv category (categorical), distance to nearest neighbor (bucketed or not), pre-print posting. Adding new ones is straightforward both conceptually and computationally, and you are then able to look at how any individual variable is correlated (via its coefficient size) to the output when taking all others into account. It may be the case for example, that once you take into account # of preprints, the effect size of distance to nearest neighbor goes to 0. What other factors can you include in your model that may be confounders?

Reviewer T5PN7/10 · confidence 3/52024-05-08

Summary

The paper analyses the use of large language models (LLMs), specifically of ChatGPT in writing scientific paper in specific fields, i.e. in sciences (Computer Science, Mathematics, Statistics, etc.). For this, they collect data from various publishers (also arXiv) and identify the usage of ChatGPT with the distributional LLM quantification. Their results show that there is an increase of ChatGPT usage in scientific writing starting from 2020 with some disciplines using it more frequently (Computer Science) than the others (Mathematics). Empiricism, Data, and Evaluation: The paper offers a strong empirical foundation, as the study uses a large collection of scientific papers (and also abstracts) and applying an established method with this data. In this way, it also offers reproducible results. Ambition, Vision, Forward-outlook: The growth in using LLMs in the recent years is very fast. The challenges posed are important to be addressed in such studies. This work is timely and highly relevant. Understanding Depth, Principled Approach: The authors show a good understanding of the approach used. They also provide detailed analysis of various aspects, such as correlation with publishing pre-prints, paper length, etc. Clarity, Honesty, and Trust: The paper under review is clearly written. However, I have some comments on the structure. The introduction contains graphs which is very unusual. It would be better to restructure the paper in the way that the introduction describes aims and motivation, as well as what comes next (paper organiser) only. The other parts should come later, e.g. graph illustrations should be part of the results. Further issues are rather minor and concern formatting. For instance, quotation marks are not properly formatted and point to one direction. The authors should also indicate the date of access to the given URLs (e.g. in references).

Rating

7

Confidence

3

Ethics flag

1

Reasons to accept

The paper is highly relevant for the topic of the conference. The results are not only interesting but can also be inspiring for further research.

Reasons to reject

I do not see any reasons to reject. However, I would recommend the authors to restructure the paper for the final version for a better readability.

Reviewer Rcc26/10 · confidence 2/52024-05-10

Summary

This paper highlights growing concerns and interest in the prevalence of AI-generated text in academic publishing, spurred by anecdotal examples and evolving editorial policies. It discusses the need for systematic analysis to understand the extent and implications of AI-modified content, introducing a framework for quantifying such modifications at scale. The study applies this framework to a large dataset of academic papers across various disciplines, revealing significant growth in AI-modified text, particularly in Computer Science. It also identifies associations between AI-modification, author behaviors like preprint posting frequency, and paper characteristics like length.

Rating

6

Confidence

2

Ethics flag

1

Reasons to accept

1. The study offers a novel contribution to the field by introducing a framework for quantifying AI-modified content in academic publishing and applying it at scale across multiple disciplines. 2. With the increasing concerns and debates surrounding the use of AI-generated content in academic publishing, the study's findings are highly relevant and timely. 3. The study opens avenues for further research into the impacts of AI-generated content on scholarly communication, knowledge dissemination, and academic discourse. Future studies could build upon the framework and findings presented, exploring additional factors influencing AI-modification trends and their broader implications for the academic community.

Reasons to reject

1. The study focuses primarily on quantifying the prevalence of AI-modified content in academic publishing without delving deeply into the potential positive aspects of AI technology in this context.

Reviewer 5uwf7/10 · confidence 4/52024-05-11

Summary

The paper investigates the increase in the use of LLM-modified content in academic publications. In particular, changes in the percentage of LLM-modified content before and after the release of ChatGPT are analyzed from the viewpoints of the fields of papers, submission rates to arXiv, similarity between papers, and length of papers. The main finding is that the proportion of LLM-modified content in computer science papers was 17.5%, followed by Electrical Engineering and Systems Science at 14.4%, which is higher than that of other fields. The authors also report that higher levels of LLM revision are associated with papers in which the first author submits preprints more frequently, papers in crowded fields, and papers that are shorter in length.

Rating

7

Confidence

4

Ethics flag

1

Reasons to accept

- The paper addresses the interesting topic of the use of LLM in academic writing. - The analysis is based on relatively large data sets, and the analytical methods are considered to have a certain degree of reliability. - While there are no major surprises about the findings, the authors have found a clear trend in the use of LLMs.

Reasons to reject

- The related work section is a weakness of this paper, as it only mentions studies related to determining whether a given sentence is LLM-modified. The authors should indicate whether there are studies that analyze the proportion of LLM use as addressed in this paper, if they exist, what they are, and how this paper stands compared to those studies. - Although there are related references in the “limitation” section, it cannot be denied that one of the reasons for the significant change in trends before and after the appearance of ChatGPT in the field of computer science is the existence of many papers on LLM and studies using LLM in this field, and the possibility that the number of contents that could be mistakenly identified as LLM-modified has increased.

Questions to authors

- In the case of journal papers, a certain amount of time is considered to have elapsed between submission and acceptance. Is the analysis based on the date of submission or publication? If it is the latter, isn't some kind of correction necessary? - Whether the collected paper set is crowded or not depends on its coverage and the breadth of the field covered. As an extreme example, the field covered by Nature is very broad, and on the other hand, it does not cover all papers in the related field, so the paper set is unlikely to be crowded. Is it reasonable to discuss whether an area is crowded or not in a situation where different sources of the set of papers are used?

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC