中文
相关论文

相关论文: CAP: Detecting Unauthorized Data Usage in Generati…

200 篇论文

Intelligent or generative writing tools rely on large language models that recognize, summarize, translate, and predict content. This position paper probes the copyright interests of open data sets used to train large language models…

计算机与社会 · 计算机科学 2023-04-07 Madiha Zahrah Choksi , David Goedicke

Exact unlearning was first introduced as a privacy mechanism that allowed a user to retract their data from machine learning models on request. Shortly after, inexact schemes were proposed to mitigate the impractical costs associated with…

Large Language Models (LLMs) have revolutionized Natural Language Processing (NLP) but pose risks of inadvertently exposing copyrighted or proprietary data, especially when such data is used for training but not intended for distribution.…

计算与语言 · 计算机科学 2025-09-16 Guangwei Zhang , Qisheng Su , Jiateng Liu , Cheng Qian , Yanzhou Pan , Yanjie Fu , Denghui Zhang

The rapid advancement of photorealistic generative models has made it increasingly important to attribute the origin of synthetic content, moving beyond binary real or fake detection toward identifying the specific model that produced a…

机器学习 · 计算机科学 2026-01-05 Ellie Thieu , Jifan Zhang , Haoyue Bai

Since its introduction in 2022, Generative AI has significantly impacted the art world, from winning state art fairs to creating complex videos from simple prompts. Amid this renaissance, a pivotal issue emerges: should users of Generative…

计算机与社会 · 计算机科学 2024-06-19 Yiyang Mei

AI-powered programming language generation (PLG) models have gained increasing attention due to their ability to generate source code of programs in a few seconds with a plain program description. Despite their remarkable performance, many…

密码学与安全 · 计算机科学 2023-05-23 Wanlun Ma , Yiliao Song , Minhui Xue , Sheng Wen , Yang Xiang

Detecting overfitting in generative models is an important challenge in machine learning. In this work, we formalize a form of overfitting that we call {\em{data-copying}} -- where the generative model memorizes and outputs training samples…

机器学习 · 计算机科学 2020-04-14 Casey Meehan , Kamalika Chaudhuri , Sanjoy Dasgupta

The growing trend of legal disputes over the unauthorized use of data in machine learning (ML) systems highlights the urgent need for reliable data-use auditing mechanisms to ensure accountability and transparency in ML. We present the…

密码学与安全 · 计算机科学 2025-09-17 Zonghao Huang , Neil Zhenqiang Gong , Michael K. Reiter

$ $The usage of generative artificial intelligence (AI) tools based on large language models, including ChatGPT, Bard, and Claude, for text generation has many exciting applications with the potential for phenomenal productivity gains. One…

计算与语言 · 计算机科学 2024-01-01 Antônio Junior Alves Caiado , Michael Hahsler

Generative Artificial Intelligence (Gen-AI) models are increasingly used to produce content across domains, including text, images, and audio. While these models represent a major technical breakthrough, they gain their generative…

机器学习 · 计算机科学 2024-12-13 Pascal Epple , Igor Shilov , Bozhidar Stevanoski , Yves-Alexandre de Montjoye

To help enforce data-protection regulations such as GDPR and detect unauthorized uses of personal data, we develop a new \emph{model auditing} technique that helps users check if their data was used to train a machine learning model. We…

密码学与安全 · 计算机科学 2019-05-21 Congzheng Song , Vitaly Shmatikov

The rapidity with which generative AI has been adopted and advanced has raised legal and ethical questions related to the impact on artists rights, content production, data collection, privacy, accuracy of information, and intellectual…

计算机与社会 · 计算机科学 2023-11-28 Cherie M Poland

Despite the rapid evolution and increasing efficacy of language and vision generative models, there remains a lack of comprehensive datasets that bridge the gap between personalized fashion needs and AI-driven design, limiting the potential…

计算机视觉与模式识别 · 计算机科学 2024-09-16 Georgia Argyrou , Angeliki Dimitriou , Maria Lymperaiou , Giorgos Filandrianos , Giorgos Stamou

The accuracy of Generative AI is increasingly critical as Large Language Models become more widely adopted. Due to potential flaws in training data and hallucination in outputs, inaccuracy can significantly impact individuals interests by…

计算机与社会 · 计算机科学 2024-07-19 Zihao Li , Weiwei Yi , Jiahong Chen

There is a growing concern that generative AI models will generate outputs closely resembling the copyrighted materials for which they are trained. This worry has intensified as the quality and complexity of generative models have immensely…

机器学习 · 计算机科学 2024-03-26 Niva Elkin-Koren , Uri Hacohen , Roi Livni , Shay Moran

The race to train language models on vast, diverse, and inconsistently documented datasets has raised pressing concerns about the legal and ethical risks for practitioners. To remedy these practices threatening data transparency and…

Deep generative models have emerged as an exciting avenue for inverse molecular design, with progress coming from the interplay between training algorithms and molecular representations. One of the key challenges in their applicability to…

Large language models (LLMs) are widely used, but concerns about data contamination challenge the reliability of LLM evaluations. Existing contamination detection methods are often task-specific or require extra prerequisites, limiting…

计算与语言 · 计算机科学 2024-10-22 Yi Zhao , Jing Li , Linyi Yang

The advent of Generative Artificial Intelligence (GenAI) models, including GitHub Copilot, OpenAI GPT, and Stable Diffusion, has revolutionized content creation, enabling non-professionals to produce high-quality content across various…

计算机视觉与模式识别 · 计算机科学 2024-05-08 Uri Hacohen , Adi Haviv , Shahar Sarfaty , Bruria Friedman , Niva Elkin-Koren , Roi Livni , Amit H Bermano

Generative Artificial Intelligence (GAI) systems that can automatically generate content in the form of source code or other contents (e.g., images) has seen increasing popularity due to the emergence of tools such as ChatGPT which rely on…

软件工程 · 计算机科学 2026-04-27 Shin Hwei Tan , Haibo Wang , Heng Li