Attribution Quality in AI-Generated Content:Benchmarking Style Embeddings and LLM Judges

Attributing authorship in the era of large language models (LLMs) is increasingly challenging as machine-generated prose rivals human writing. We benchmark two complementary attribution mechanisms—fixed Style Embeddings and an instruction-tuned LLM judge (GPT-40)—on the Human-AI Parallel Corpus, an open dataset from which we choose 600 balanced instances spanning six domains (academic, news, fiction, blogs, spoken transcripts, and TV/movie scripts). Each instance contains a human prompt with both a gold continuation and an LLM-generated continuation from either GPT-40 or LLAMA-70B-Instruct. The Style Embedding baseline achieves stronger aggregate accuracy on GPT continuations (82 % vs. 68 %). The LLM Judge is slightly better than the Style embeddings on LLaMA continuations (85 % vs. 81 %) but the results are not statistically significant. Crucially, the LLM judge significantly outperforms in fiction and academic prose, indicating semantic sensitivity, whereas embeddings dominate in spoken and scripted dialogue, reflecting structural strengths. These complementary patterns highlight attribution as a multidimensional problem requiring hybrid strategies. To support reproducibility we provide code in GitHub repositories and derived data on Huggingface. We release both the dataset and source code under the MIT license. This open framework provides a reproducible benchmark for attribution quality assessment in AI-generated content. We also provide a thorough review of prior literature particularly the papers that directly influenced the framework and results in this paper.

Paper

References (21)

Scroll for more · 9 remaining

Similar papers

© 2026 NYSGPT2525 LLC