English
Related papers

Related papers: Blameless Users in a Clean Room: Defining Copyrigh…

200 papers

There is a growing concern that learned conditional generative models may output samples that are substantially similar to some copyrighted data $C$ that was in their training set. We give a formal definition of $\textit{near…

Machine Learning · Computer Science 2023-07-24 Nikhil Vyas , Sham Kakade , Boaz Barak

To achieve accurate and unbiased predictions, Machine Learning (ML) models rely on large, heterogeneous, and high-quality datasets. However, this could raise ethical and legal concerns regarding copyright and authorization aspects,…

Machine Learning · Computer Science 2024-10-10 Daniela Gallo , Angelica Liguori , Ettore Ritacco , Luca Caviglione , Fabrizio Durante , Giuseppe Manco

Generative AI has witnessed rapid advancement in recent years, expanding their capabilities to create synthesized content such as text, images, audio, and code. The high fidelity and authenticity of contents generated by these Deep…

Cryptography and Security · Computer Science 2024-07-25 Jie Ren , Han Xu , Pengfei He , Yingqian Cui , Shenglai Zeng , Jiankun Zhang , Hongzhi Wen , Jiayuan Ding , Pei Huang , Lingjuan Lyu , Hui Liu , Yi Chang , Jiliang Tang

In this paper, we investigate potential randomization approaches that can complement current practices of input-based methods (such as licensing data and prompt filtering) and output-based methods (such as recitation checker, license…

Cryptography and Security · Computer Science 2024-08-27 Wei-Ning Chen , Peter Kairouz , Sewoong Oh , Zheng Xu

There is a growing concern that generative AI models will generate outputs closely resembling the copyrighted materials for which they are trained. This worry has intensified as the quality and complexity of generative models have immensely…

Machine Learning · Computer Science 2024-03-26 Niva Elkin-Koren , Uri Hacohen , Roi Livni , Shay Moran

Copyright infringement may occur when a generative model produces samples substantially similar to some copyrighted data that it had access to during the training phase. The notion of access usually refers to including copyrighted samples…

Machine Learning · Computer Science 2024-06-05 Yiwei Lu , Matthew Y. R. Yang , Zuoqiu Liu , Gautam Kamath , Yaoliang Yu

As AI advances, copyrighted content faces growing risk of unauthorized use, whether through model training or direct misuse. Building upon invisible adversarial perturbation, recent works developed copyright protections against specific AI…

Machine Learning · Computer Science 2025-06-04 Tianci Liu , Tong Yang , Quan Zhang , Qi Lei

This paper presents a probabilistic approach to analyzing copyright infringement disputes. Evidentiary principles shaped by case law are formalized in probabilistic terms, and the ``inverse ratio rule'' -- a controversial legal doctrine…

Computers and Society · Computer Science 2026-01-21 Hiroaki Chiba-Okabe

The risk of language models unintentionally reproducing copyrighted material from their training data has led to the development of various protective measures. In this paper, we propose model fusion as an effective solution to safeguard…

Machine Learning · Computer Science 2024-07-30 Javier Abad , Konstantin Donhauser , Francesco Pinto , Fanny Yang

Studying data memorization in neural language models helps us understand the risks (e.g., to privacy or copyright) associated with models regurgitating training data and aids in the development of countermeasures. Many prior works -- and…

Large language models (LLMs) commonly risk copyright infringement by reproducing protected content verbatim or with insufficient transformative modifications, posing significant ethical, legal, and practical concerns. Current inference-time…

Computation and Language · Computer Science 2025-06-02 Aakash Sen Sharma , Debdeep Sanyal , Priyansh Srivastava , Sundar Atreya H. , Shirish Karande , Mohan Kankanhalli , Murari Mandal

Generative models have achieved impressive results in text to image tasks, significantly advancing visual content creation. However, this progress comes at a cost, as such models rely heavily on large-scale training data and may…

Machine Learning · Computer Science 2025-09-03 Zhipeng Yin , Zichong Wang , Avash Palikhe , Zhen Liu , Jun Liu , Wenbin Zhang

Copyright law focuses on whether a new work is "substantially similar" to an existing one, but generative AI can closely imitate style without copying content, a capability now central to ongoing litigation. We argue that existing…

Theoretical Economics · Economics 2026-02-13 Annie Liang , Jay Lu

In this paper, we highlight a critical threat posed by emerging neural models: data plagiarism. We demonstrate how modern neural models (e.g., diffusion models) can replicate copyrighted images, even when protected by advanced watermarking…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zihang Zou , Boqing Gong , Liqiang Wang

Diffusion models have demonstrated remarkable performance in image generation tasks, paving the way for powerful AIGC applications. However, these widely-used generative models can also raise security and privacy concerns, such as copyright…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Zhengyue Zhao , Jinhao Duan , Xing Hu , Kaidi Xu , Chenan Wang , Rui Zhang , Zidong Du , Qi Guo , Yunji Chen

AI-based colorization has shown remarkable capability in generating realistic color images from grayscale inputs. However, it poses risks of copyright infringement -- for example, the unauthorized colorization and resale of monochrome manga…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Yuki Nii , Futa Waseda , Ching-Chun Chang , Isao Echizen

The widespread use of Large Language Models (LLMs) raises critical concerns regarding the unauthorized inclusion of copyrighted content in training data. Existing detection frameworks, such as DE-COP, are computationally intensive, and…

Artificial Intelligence · Computer Science 2026-03-20 David Szczecina , Senan Gaffori , Edmond Li

Language models may memorize more than just facts, including entire chunks of texts seen during training. Fair use exemptions to copyright laws typically allow for limited use of copyrighted material without permission from the copyright…

Computation and Language · Computer Science 2023-10-24 Antonia Karamolegkou , Jiaang Li , Li Zhou , Anders Søgaard

Training a deep neural network (DNN) requires a high computational cost. Buying models from sellers with a large number of computing resources has become prevailing. However, the buyer-seller environment is not always trusted. To protect…

Cryptography and Security · Computer Science 2023-12-12 Yusheng Guo , Nan Zhong , Zhenxing Qian , Xinpeng Zhang

Generative AI (e.g., Generative Adversarial Networks - GANs) has become increasingly popular in recent years. However, Generative AI introduces significant concerns regarding the protection of Intellectual Property Rights (IPR) (resp. model…

‹ Prev 1 2 3 10 Next ›