English
Related papers

Related papers: Dataset Ownership Verification for Pre-trained Mas…

200 papers

Deep learning has achieved remarkable progress in various applications, heightening the importance of safeguarding the intellectual property (IP) of well-trained models. It entails not only authorizing usage but also ensuring the deployment…

Computer Vision and Pattern Recognition · Computer Science 2024-03-08 Boyang Peng , Sanqing Qu , Yong Wu , Tianpei Zou , Lianghua He , Alois Knoll , Guang Chen , changjun jiang

Large Language Models (LLMs) have shown their impressive capabilities, while also raising concerns about the data contamination problems due to privacy issues and leakage of benchmark datasets in the pre-training phase. Therefore, it is…

Computation and Language · Computer Science 2024-06-04 Zhenhua Liu , Tong Zhu , Chuanyuan Tan , Haonan Lu , Bing Liu , Wenliang Chen

Protecting the use of audio datasets is a major concern for data owners, particularly with the recent rise of audio deep learning models. While watermarks can be used to protect the data itself, they do not allow to identify a deep learning…

Cryptography and Security · Computer Science 2025-03-14 Wassim Bouaziz , El-Mahdi El-Mhamdi , Nicolas Usunier

The huge supporting training data on the Internet has been a key factor in the success of deep learning models. However, this abundance of public-available data also raises concerns about the unauthorized exploitation of datasets for…

Cryptography and Security · Computer Science 2023-04-11 Ruixiang Tang , Qizhang Feng , Ninghao Liu , Fan Yang , Xia Hu

Training data is a critical and often proprietary asset in Large Language Model (LLM) development, motivating the use of data watermarking to embed model-transferable signals for usage verification. We identify low coverage as a vital yet…

Cryptography and Security · Computer Science 2026-04-30 Hengyu Wu , Yang Cao

Self-supervised models are increasingly prevalent in machine learning (ML) since they reduce the need for expensively labeled data. Because of their versatility in downstream applications, they are increasingly used as a service exposed via…

Auditing the use of data in training machine-learning (ML) models is an increasingly pressing challenge, as myriad ML practitioners routinely leverage the effort of content creators to train models without their permission. In this paper,…

Cryptography and Security · Computer Science 2025-01-28 Zonghao Huang , Neil Zhenqiang Gong , Michael K. Reiter

Recent years have witnessed tremendous success in Self-Supervised Learning (SSL), which has been widely utilized to facilitate various downstream tasks in Computer Vision (CV) and Natural Language Processing (NLP) domains. However,…

Cryptography and Security · Computer Science 2024-01-30 Peizhuo Lv , Pan Li , Shenchen Zhu , Shengzhi Zhang , Kai Chen , Ruigang Liang , Chang Yue , Fan Xiang , Yuling Cai , Hualong Ma , Yingjun Zhang , Guozhu Meng

Diffusion Models (DMs) benefit from large and diverse datasets for their training. Since this data is often scraped from the Internet without permission from the data owners, this raises concerns about copyright and intellectual property…

Machine Learning · Computer Science 2025-06-24 Jan Dubiński , Antoni Kowalczuk , Franziska Boenisch , Adam Dziedzic

The surge in popularity of machine learning (ML) has driven significant investments in training Deep Neural Networks (DNNs). However, these models that require resource-intensive training are vulnerable to theft and unauthorized use. This…

Cryptography and Security · Computer Science 2024-03-12 Jasper Stang , Torsten Krauß , Alexandra Dmitrienko

The growing trend of legal disputes over the unauthorized use of data in machine learning (ML) systems highlights the urgent need for reliable data-use auditing mechanisms to ensure accountability and transparency in ML. We present the…

Cryptography and Security · Computer Science 2025-09-17 Zonghao Huang , Neil Zhenqiang Gong , Michael K. Reiter

The great economic values of deep neural networks (DNNs) urge AI enterprises to protect their intellectual property (IP) for these models. Recently, proof-of-training (PoT) has been proposed as a promising solution to DNN IP protection,…

Cryptography and Security · Computer Science 2024-10-11 Yijia Chang , Hanrui Jiang , Chao Lin , Xinyi Huang , Jian Weng

Text-to-image (T2I) diffusion models enable high-quality image generation conditioned on textual prompts. However, fine-tuning these pre-trained models for personalization raises concerns about unauthorized dataset usage. To address this…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Kuofeng Gao , Yufei Zhu , Yiming Li , Jiawang Bai , Yong Yang , Zhifeng Li , Shu-Tao Xia

Large language models (LLMs) are pre-trained and post-trained on vast amounts of loosely curated data, raising the possibility that these models may have been trained on proprietary datasets or the same benchmarks used for evaluation. This…

Machine Learning · Computer Science 2026-05-11 Pengrun Huang , Kamalika Chaudhuri , Yu-Xiang Wang

Machine Learning as a Service (MLaaS) has emerged as a widely adopted paradigm for providing access to deep neural network (DNN) models, enabling users to conveniently leverage these models through standardized APIs. However, such services…

Machine Learning · Computer Science 2026-02-25 Bolin Shen , Zhan Cheng , Neil Zhenqiang Gong , Fan Yao , Yushun Dong

The surging demand for large-scale datasets in deep learning has heightened the need for effective copyright protection, given the risks of unauthorized use to data owners. Although the dataset watermark technique holds promise for auditing…

Cryptography and Security · Computer Science 2026-02-17 Xiao Ren , Xinyi Yu , Linkang Du , Min Chen , Yuanchao Shu , Zhou Su , Yunjun Gao , Zhikun Zhang

Recent advancements in Deep Neural Network (DNN) models have significantly improved performance across computer vision tasks. However, achieving highly generalizable and high-performing vision models requires expansive datasets, resulting…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Minhyun Lee , Song Park , Byeongho Heo , Dongyoon Han , Hyunjung Shim

Deep learning models are usually black boxes when deployed on machine learning platforms. Prior works have shown that the attributes ($e.g.$, the number of convolutional layers) of a target black-box neural network can be exposed through a…

Machine Learning · Computer Science 2023-07-21 Rongqing Li , Jiaqi Yu , Changsheng Li , Wenhan Luo , Ye Yuan , Guoren Wang

The rapid advancement of Large Vision-Language Models (LVLMs) is increasingly accompanied by unauthorized scraping and training on multimodal web data, posing severe copyright and privacy risks to data owners. Existing countermeasures, such…

Cryptography and Security · Computer Science 2026-05-15 Chengshuai Zhao , Zhen Tan , Dawei Li , Zhiyuan Yu , Huan Liu

Deep neural networks (DNNs) have achieved tremendous success in artificial intelligence (AI) fields. However, DNN models can be easily illegally copied, redistributed, or abused by criminals, seriously damaging the interests of model…

Cryptography and Security · Computer Science 2023-11-29 Xuefeng Fan , Dahao Fu , Hangyu Gui , Xinpeng Zhang , Xiaoyi Zhou