中文
相关论文

相关论文: Dataset Ownership Verification for Pre-trained Mas…

200 篇论文

Deep learning has achieved remarkable progress in various applications, heightening the importance of safeguarding the intellectual property (IP) of well-trained models. It entails not only authorizing usage but also ensuring the deployment…

计算机视觉与模式识别 · 计算机科学 2024-03-08 Boyang Peng , Sanqing Qu , Yong Wu , Tianpei Zou , Lianghua He , Alois Knoll , Guang Chen , changjun jiang

Large Language Models (LLMs) have shown their impressive capabilities, while also raising concerns about the data contamination problems due to privacy issues and leakage of benchmark datasets in the pre-training phase. Therefore, it is…

计算与语言 · 计算机科学 2024-06-04 Zhenhua Liu , Tong Zhu , Chuanyuan Tan , Haonan Lu , Bing Liu , Wenliang Chen

Protecting the use of audio datasets is a major concern for data owners, particularly with the recent rise of audio deep learning models. While watermarks can be used to protect the data itself, they do not allow to identify a deep learning…

密码学与安全 · 计算机科学 2025-03-14 Wassim Bouaziz , El-Mahdi El-Mhamdi , Nicolas Usunier

The huge supporting training data on the Internet has been a key factor in the success of deep learning models. However, this abundance of public-available data also raises concerns about the unauthorized exploitation of datasets for…

密码学与安全 · 计算机科学 2023-04-11 Ruixiang Tang , Qizhang Feng , Ninghao Liu , Fan Yang , Xia Hu

Training data is a critical and often proprietary asset in Large Language Model (LLM) development, motivating the use of data watermarking to embed model-transferable signals for usage verification. We identify low coverage as a vital yet…

密码学与安全 · 计算机科学 2026-04-30 Hengyu Wu , Yang Cao

Self-supervised models are increasingly prevalent in machine learning (ML) since they reduce the need for expensively labeled data. Because of their versatility in downstream applications, they are increasingly used as a service exposed via…

Auditing the use of data in training machine-learning (ML) models is an increasingly pressing challenge, as myriad ML practitioners routinely leverage the effort of content creators to train models without their permission. In this paper,…

密码学与安全 · 计算机科学 2025-01-28 Zonghao Huang , Neil Zhenqiang Gong , Michael K. Reiter

Recent years have witnessed tremendous success in Self-Supervised Learning (SSL), which has been widely utilized to facilitate various downstream tasks in Computer Vision (CV) and Natural Language Processing (NLP) domains. However,…

密码学与安全 · 计算机科学 2024-01-30 Peizhuo Lv , Pan Li , Shenchen Zhu , Shengzhi Zhang , Kai Chen , Ruigang Liang , Chang Yue , Fan Xiang , Yuling Cai , Hualong Ma , Yingjun Zhang , Guozhu Meng

Diffusion Models (DMs) benefit from large and diverse datasets for their training. Since this data is often scraped from the Internet without permission from the data owners, this raises concerns about copyright and intellectual property…

机器学习 · 计算机科学 2025-06-24 Jan Dubiński , Antoni Kowalczuk , Franziska Boenisch , Adam Dziedzic

The surge in popularity of machine learning (ML) has driven significant investments in training Deep Neural Networks (DNNs). However, these models that require resource-intensive training are vulnerable to theft and unauthorized use. This…

密码学与安全 · 计算机科学 2024-03-12 Jasper Stang , Torsten Krauß , Alexandra Dmitrienko

The growing trend of legal disputes over the unauthorized use of data in machine learning (ML) systems highlights the urgent need for reliable data-use auditing mechanisms to ensure accountability and transparency in ML. We present the…

密码学与安全 · 计算机科学 2025-09-17 Zonghao Huang , Neil Zhenqiang Gong , Michael K. Reiter

The great economic values of deep neural networks (DNNs) urge AI enterprises to protect their intellectual property (IP) for these models. Recently, proof-of-training (PoT) has been proposed as a promising solution to DNN IP protection,…

密码学与安全 · 计算机科学 2024-10-11 Yijia Chang , Hanrui Jiang , Chao Lin , Xinyi Huang , Jian Weng

Text-to-image (T2I) diffusion models enable high-quality image generation conditioned on textual prompts. However, fine-tuning these pre-trained models for personalization raises concerns about unauthorized dataset usage. To address this…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Kuofeng Gao , Yufei Zhu , Yiming Li , Jiawang Bai , Yong Yang , Zhifeng Li , Shu-Tao Xia

Large language models (LLMs) are pre-trained and post-trained on vast amounts of loosely curated data, raising the possibility that these models may have been trained on proprietary datasets or the same benchmarks used for evaluation. This…

机器学习 · 计算机科学 2026-05-11 Pengrun Huang , Kamalika Chaudhuri , Yu-Xiang Wang

Machine Learning as a Service (MLaaS) has emerged as a widely adopted paradigm for providing access to deep neural network (DNN) models, enabling users to conveniently leverage these models through standardized APIs. However, such services…

机器学习 · 计算机科学 2026-02-25 Bolin Shen , Zhan Cheng , Neil Zhenqiang Gong , Fan Yao , Yushun Dong

The surging demand for large-scale datasets in deep learning has heightened the need for effective copyright protection, given the risks of unauthorized use to data owners. Although the dataset watermark technique holds promise for auditing…

密码学与安全 · 计算机科学 2026-02-17 Xiao Ren , Xinyi Yu , Linkang Du , Min Chen , Yuanchao Shu , Zhou Su , Yunjun Gao , Zhikun Zhang

Recent advancements in Deep Neural Network (DNN) models have significantly improved performance across computer vision tasks. However, achieving highly generalizable and high-performing vision models requires expansive datasets, resulting…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Minhyun Lee , Song Park , Byeongho Heo , Dongyoon Han , Hyunjung Shim

Deep learning models are usually black boxes when deployed on machine learning platforms. Prior works have shown that the attributes ($e.g.$, the number of convolutional layers) of a target black-box neural network can be exposed through a…

机器学习 · 计算机科学 2023-07-21 Rongqing Li , Jiaqi Yu , Changsheng Li , Wenhan Luo , Ye Yuan , Guoren Wang

The rapid advancement of Large Vision-Language Models (LVLMs) is increasingly accompanied by unauthorized scraping and training on multimodal web data, posing severe copyright and privacy risks to data owners. Existing countermeasures, such…

密码学与安全 · 计算机科学 2026-05-15 Chengshuai Zhao , Zhen Tan , Dawei Li , Zhiyuan Yu , Huan Liu

Deep neural networks (DNNs) have achieved tremendous success in artificial intelligence (AI) fields. However, DNN models can be easily illegally copied, redistributed, or abused by criminals, seriously damaging the interests of model…

密码学与安全 · 计算机科学 2023-11-29 Xuefeng Fan , Dahao Fu , Hangyu Gui , Xinpeng Zhang , Xiaoyi Zhou