Measuring Copyright Risks of Large Language Model via Partial Information Probing

Exploring the data sources used to train Large Language Models (LLMs) is a crucial direction in investigating potential copyright infringement by these models. While this approach can identify the possible use of copyrighted materials in training data, it does not directly measure infringing risks. Recent research has shifted towards testing whether LLMs can directly output copyrighted content. Addressing this direction, we investigate and assess LLMs' capacity to generate infringing content by providing them with partial information from copyrighted materials, and try to use iterative prompting to get LLMs to generate more infringing content. Specifically, we input a portion of a copyrighted text into LLMs, prompt them to complete it, and then analyze the overlap between the generated content and the original copyrighted material. Our findings demonstrate that LLMs can indeed generate content highly overlapping with copyrighted materials based on these partial inputs.

Paper

References (16)

06Automaticevaluationofmachinetrans-lation quality using longest common subsequence and skip-bigramstatistics2004
07CopyrightLaw of the United States and Related Laws Contained in Tıtle 17 of the United States Code1976
082024. A survey on large language model (LLM) security and privacy: The Good, The Bad, and The UglyHigh-Confidence Computing
092024. LLMs and Memorization: On Quality and Specificity of Copyright CompliancearXiv
10DCAI24, Boise, ID, United StatesProceedings of the 42nd annual meeting of the association for computational linguistics (ACL-04)
112023. Beyond fair use: Legal risk evaluation for training LLMs on copyrighted textICML Workshop on Gener-ative AI and Law
122024. The role of llms in sustainable smart cities: Applications, challenges, and future directionsarXiv preprint

Scroll for more · 4 remaining

Similar papers

© 2026 NYSGPT2525 LLC