AI-generated text corpus AI-GenT 1.0

The AI-Generated Text (AI-GenT) corpus is a collection of English and Slovenian texts generated by several large language models. The corpus has been used in comparisons to collections of human-written texts in order to investigate the linguistic characteristics of the language generated by LLMs. The current version of the corpus contains texts that were constructed based on two preexisting human-written text corpora: the Šolar 3.0 corpus of Slovenian student essays (http://hdl.handle.net/11356/1589) and the LOCNESS corpus of English native speaker student essays (provided by the Centre for English Corpus Linguistics (CECL) at Université catholique de Louvain in Belgium - https://www.learnercorpusassociation.org/resources/tools/locness-corpus/). Three different LLMs—GPT-5 (https://developers.openai.com/api/docs/models/gpt-5), GaMS-27B (https://huggingface.co/cjvt/GaMS-27B-Instruct), and gemma-2-27b (https://huggingface.co/google/gemma-2-27b-it)—were instructed to produce corresponding texts to the texts in the human-written corpora using prompts containing information about the topic and length of the desired output. The AI-generated texts were produced by taking various subsets of the original human-written corpora as the basis for constructing the input prompts. For a full overview of the data, model, and prompt type combinations used to generate the AI-generated texts, please refer to the included AI-GenT_structure.png file which includes a full visual representation of the corpus structure. The corpus contains the AI-generated texts both in the form of raw text files as well as in the CoNLL-U file format containing grammatical annotations following the UD system of annotation (https://universaldependencies.org/). UD annotations were generated using the Trankit NLP pipeline (https://aclanthology.org/2021.eacl-demos.10/) with the default model used for English and a custom model used for Slovenian that is retrained on UD v2.15 data (http://hdl.handle.net/11356/1997). In the future, the corpus is planned to be extended with additional AI-generated news articles and Wikipedia articles.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC