When the Instrument Studies Itself: A Systematic Map of Autonomous AI Research Systems (2024-2026), Conducted by a Frontier Language Model
A systematic map of 811 papers on autonomous AI research systems (January 2024 - July 2026), conducted by a frontier language model, with its own research process logged and analysed as data. AUTHORSHIP. This work was jointly produced by a human researcher and an AI system. Claude Fable 5 (Anthropic) proposed the research direction, wrote the review protocol and codebook, wrote all software, executed the search, screening, coding and analysis, generated the figures, and drafted the manuscript. Khristian Kopachelli framed the project, made the scope decisions, overruled the AI's judgment where he disagreed (one such reversal changed the corpus by 183 papers), contributed the two conceptual arguments in the Discussion, and takes responsibility for the published work. This record lists both as creators. The companion arXiv version lists only the human author, because arXiv policy does not permit AI systems to be named as authors; the two records cross-reference each other and describe the same collaboration. FINDINGS. Quarterly output in this literature grew roughly fortyfold in nine quarters. 71% of papers report benchmark metrics, but only 5% re-test a result outside the loop that produced it, 55% provide no auditability mechanism of any kind, and 71% never state what humans did. Conditioning on claim strength reveals a dissociation: papers claiming the AI produced new scientific knowledge validate their RESULTS better than average (50% report external validation versus 11% corpus-wide) while exposing their PROCESS worse (64% provide no auditability mechanism versus 55%). This field validates outputs better than it exposes process, and the gap is widest where its claims are strongest. REFLEXIVE COMPONENT. Because the review was conducted by the class of system it studies, the process was logged rather than smoothed. Two frontier models screening the same corpus under the same protocol agreed on 83.3% of decisions (Cohen's kappa 0.71), with disagreements clustering on seven boundary classes where the field has not settled what counts as an AI scientist, forcing a documented protocol amendment mid-review. Six AI failures are reported with measured rates, including a workflow that reported success having performed no work, silent item-dropping in structured extraction at 0.4-0.5%, a pipeline defect that fed empty inputs to a reasoning step, and a draft thesis contradicted by the authors' own data. CONTENTS. Paper (PDF), full candidate corpus and both layers of screening decisions with per-paper justifications, the coded corpus, mechanical full-text artifact-verification results for all 811 papers, all analysis and figure code, the review protocol with amendments, and the three contemporaneous process logs. Code is released under the MIT License; data, text and figures under CC BY 4.0. A NOTE ON PROVENANCE. Brzozowski and Chung (arXiv:2606.02184) document 1,655 Zenodo records bearing real DOIs, fabricated authors, and backdated timestamps. This deposit is intended as the inverse in every respect: an accountable human author, an AI contributor named as exactly what it is, unmanipulated timestamps, released underlying data, and a complete log of how the work was produced, including its failures.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex