中文
相关论文

相关论文: CAP: Detecting Unauthorized Data Usage in Generati…

200 篇论文

Large Generative AI (GAI) models have the unparalleled ability to generate text, images, audio, and other forms of media that are increasingly indistinguishable from human-generated content. As these models often train on publicly available…

计算机与社会 · 计算机科学 2024-06-25 Tanja Šarčević , Alicja Karlowicz , Rudolf Mayer , Ricardo Baeza-Yates , Andreas Rauber

The widespread use of Large Language Models (LLMs) raises critical concerns regarding the unauthorized inclusion of copyrighted content in training data. Existing detection frameworks, such as DE-COP, are computationally intensive, and…

人工智能 · 计算机科学 2026-03-20 David Szczecina , Senan Gaffori , Edmond Li

Exploring the data sources used to train Large Language Models (LLMs) is a crucial direction in investigating potential copyright infringement by these models. While this approach can identify the possible use of copyrighted materials in…

计算与语言 · 计算机科学 2024-09-24 Weijie Zhao , Huajie Shao , Zhaozhuo Xu , Suzhen Duan , Denghui Zhang

Generative AI has witnessed rapid advancement in recent years, expanding their capabilities to create synthesized content such as text, images, audio, and code. The high fidelity and authenticity of contents generated by these Deep…

Generative AI is becoming increasingly prevalent in creative fields, sparking urgent debates over how current copyright laws can keep pace with technological innovation. Recent controversies of AI models generating near-replicas of…

机器学习 · 计算机科学 2025-07-01 Archer Amon , Zhipeng Yin , Zichong Wang , Avash Palikhe , Wenbin Zhang

There is a growing concern that learned conditional generative models may output samples that are substantially similar to some copyrighted data $C$ that was in their training set. We give a formal definition of $\textit{near…

机器学习 · 计算机科学 2023-07-24 Nikhil Vyas , Sham Kakade , Boaz Barak

Generative models have achieved impressive results in text to image tasks, significantly advancing visual content creation. However, this progress comes at a cost, as such models rely heavily on large-scale training data and may…

机器学习 · 计算机科学 2025-09-03 Zhipeng Yin , Zichong Wang , Avash Palikhe , Zhen Liu , Jun Liu , Wenbin Zhang

Text-To-Music (TTM) models have recently revolutionized the automatic music generation research field. Specifically, by reaching superior performances to all previous state-of-the-art models and by lowering the technical proficiency needed…

音频与语音处理 · 电气工程与系统科学 2024-09-26 Luca Comanducci , Paolo Bestagini , Stefano Tubaro

Questions of fair use of copyright-protected content to train Large Language Models (LLMs) are being actively debated. Document-level inference has been proposed as a new task: inferring from black-box access to the trained model whether a…

计算与语言 · 计算机科学 2024-06-06 Matthieu Meeus , Igor Shilov , Manuel Faysse , Yves-Alexandre de Montjoye

As machine- and AI-generated content proliferates, protecting the intellectual property of generative models has become imperative, yet verifying data ownership poses formidable challenges, particularly in cases of unauthorized reuse of…

机器学习 · 计算机科学 2024-02-28 Aditya Desu , Xuanli He , Qiongkai Xu , Wei Lu

Copyright infringement may occur when a generative model produces samples substantially similar to some copyrighted data that it had access to during the training phase. The notion of access usually refers to including copyrighted samples…

机器学习 · 计算机科学 2024-06-05 Yiwei Lu , Matthew Y. R. Yang , Zuoqiu Liu , Gautam Kamath , Yaoliang Yu

Generative AI models, renowned for their ability to synthesize high-quality content, have sparked growing concerns over the improper generation of copyright-protected material. While recent studies have proposed various approaches to…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Qipan Xu , Zhenting Wang , Xiaoxiao He , Ligong Han , Ruixiang Tang

Generative artificial intelligence (AI) systems are trained on large data corpora to generate new pieces of text, images, videos, and other media. There is growing concern that such systems may infringe on the copyright interests of…

机器学习 · 计算机科学 2024-09-10 Jiachen T. Wang , Zhun Deng , Hiroaki Chiba-Okabe , Boaz Barak , Weijie J. Su

The advent of generative AI models has revolutionized digital content creation, yet it introduces challenges in maintaining copyright integrity due to generative parroting, where models mimic their training data too closely. Our research…

机器学习 · 计算机科学 2024-06-21 Saeid Asgari Taghanaki , Joseph Lambourne

Large language models (LLMs) trained on unfiltered corpora inherently risk retaining sensitive information, necessitating selective knowledge unlearning for regulatory compliance and ethical safety. However, existing parameter-modifying…

机器学习 · 计算机科学 2026-05-18 Zhaokun Wang , Jinyu Guo , Jingwen Pu , Hongli Pu , Meng Yang , Xunlei Chen , Jie Ou , Wenyi Li , Guangchun Luo , Wenhong Tian

Are there any conditions under which a generative model's outputs are guaranteed not to infringe the copyrights of its training data? This is the question of "provable copyright protection" first posed by Vyas, Kakade, and Barak (ICML…

密码学与安全 · 计算机科学 2026-02-26 Aloni Cohen

Large scale text-to-image generation models can memorize and reproduce their training dataset. Since the training dataset often contains copyrighted material, reproduction of training dataset poses a copyright infringement risk, which could…

机器学习 · 计算机科学 2025-12-18 Neeraj Sarna , Yuanyuan Li , Michael von Gablenz

The rapid progress of generative AI technology has sparked significant copyright concerns, leading to numerous lawsuits filed against AI developers. Notably, generative AI's capacity for generating images of copyrighted characters has been…

机器学习 · 计算机科学 2025-04-01 Hiroaki Chiba-Okabe , Weijie J. Su

Content is created for a well-defined purpose, often described by a metric or signal represented in the form of structured information. The relationship between the goal (metrics) of target content and the content itself is non-trivial.…

计算与语言 · 计算机科学 2022-03-29 Navita Goyal , Roodram Paneri , Ayush Agarwal , Udit Kalani , Abhilasha Sancheti , Niyati Chhaya

As generative models become powerful, concerns around transparency, accountability, and copyright violations have intensified. Understanding how specific training data contributes to a model's output is critical. We introduce a framework…

人工智能 · 计算机科学 2025-12-03 Theodoros Aivalis , Iraklis A. Klampanos , Antonis Troumpoukis , Joemon M. Jose
‹ 上一页 1 2 3 10 下一页 ›