中文
相关论文

相关论文: Provable Watermarking for Data Poisoning Attacks

200 篇论文

In this paper, we introduce a simple yet effective tabular data watermarking mechanism with statistical guarantees. We show theoretically that the proposed watermark can be effectively detected, while faithfully preserving the data…

密码学与安全 · 计算机科学 2024-05-28 Hengzhi He , Peiyu Yu , Junpeng Ren , Ying Nian Wu , Guang Cheng

Constructing and curating high-quality code datasets requires significant resources, making them valuable intellectual property. Unfortunately, these datasets currently face severe risks of unauthorized use. Although digital watermarking…

软件工程 · 计算机科学 2026-05-01 Haocheng Huang , Yuchen Chen , Weisong Sun , Peizhuo Lv , Yuan Xiao , Chunrong Fang , Yang Liu , Xiaofang Zhang

Watermarking has emerged as an effective solution for copyright protection of synthetic data. However, applying watermarking techniques to synthetic tabular data presents challenges, as tabular data can easily lose their watermarks through…

密码学与安全 · 计算机科学 2026-03-17 Yuyang Xia , Yaoqiang Xu , Chen Qian , Yang Li , Guoliang Li , Jianhua Feng

In this paper, we analyze several recent schemes for watermarking network flows that are based on splitting the flow into timing intervals. We show that this approach creates time-dependent correlations that enable an attack that combines…

密码学与安全 · 计算机科学 2015-03-20 Negar Kiyavash , Amir Houmansadr , Nikita Borisov

The promise of LLM watermarking rests on a core assumption that a specific watermark proves authorship by a specific model. We demonstrate that this assumption is dangerously flawed. We introduce the threat of watermark spoofing, a…

密码学与安全 · 计算机科学 2026-02-24 Hyeseon An , Shinwoo Park , Suyeon Woo , Yo-Sub Han

Data poisoning attacks aim to manipulate the model produced by a learning algorithm by adversarially modifying the training set. We consider differential privacy as a defensive measure against this type of attack. We show that such learners…

机器学习 · 计算机科学 2019-07-08 Yuzhe Ma , Xiaojin Zhu , Justin Hsu

Data poisoning -- the process by which an attacker takes control of a model by making imperceptible changes to a subset of the training data -- is an emerging threat in the context of neural networks. Existing attacks for data poisoning…

机器学习 · 计算机科学 2021-02-23 W. Ronny Huang , Jonas Geiping , Liam Fowl , Gavin Taylor , Tom Goldstein

As deep learning (DL) models are widely and effectively used in Machine Learning as a Service (MLaaS) platforms, there is a rapidly growing interest in DL watermarking techniques that can be used to confirm the ownership of a particular…

密码学与安全 · 计算机科学 2024-11-22 Mikhail Pautov , Nikita Bogdanov , Stanislav Pyatkin , Oleg Rogov , Ivan Oseledets

Recently, point clouds have been widely used in computer vision, whereas their collection is time-consuming and expensive. As such, point cloud datasets are the valuable intellectual property of their owners and deserve protection. To…

密码学与安全 · 计算机科学 2024-11-05 Cheng Wei , Yang Wang , Kuofeng Gao , Shuo Shao , Yiming Li , Zhibo Wang , Zhan Qin

Data Poisoning attacks modify training data to maliciously control a model trained on such data. In this work, we focus on targeted poisoning attacks which cause a reclassification of an unmodified test image and as such breach model…

计算机视觉与模式识别 · 计算机科学 2021-05-11 Jonas Geiping , Liam Fowl , W. Ronny Huang , Wojciech Czaja , Gavin Taylor , Michael Moeller , Tom Goldstein

Protecting the use of audio datasets is a major concern for data owners, particularly with the recent rise of audio deep learning models. While watermarks can be used to protect the data itself, they do not allow to identify a deep learning…

密码学与安全 · 计算机科学 2025-03-14 Wassim Bouaziz , El-Mahdi El-Mhamdi , Nicolas Usunier

Data watermarking in language models injects traceable signals, such as specific token sequences or stylistic patterns, into copyrighted text, allowing copyright holders to track and verify training data ownership. Previous data…

密码学与安全 · 计算机科学 2025-07-29 Xinyue Cui , Johnny Tian-Zheng Wei , Swabha Swayamdipta , Robin Jia

Recent years have seen a surge in interest in digital content watermarking techniques, driven by the proliferation of generative models and increased legal pressure. With an ever-growing percentage of AI-generated content available online,…

The rapid advancement of deep neural networks (DNNs) heavily relies on large-scale, high-quality datasets. However, unauthorized commercial use of these datasets severely violates the intellectual property rights of dataset owners. Existing…

密码学与安全 · 计算机科学 2025-10-31 Yingjia Wang , Ting Qiao , Xing Liu , Chongzuo Li , Sixing Wu , Jianbin Li

Watermarking enables GenAI providers to verify whether content was generated by their models. A watermark is a hidden signal in the content, whose presence can be detected using a secret watermark key. A core security threat are forgery…

密码学与安全 · 计算机科学 2026-05-12 Toluwani Aremu , Noor Hussein , Munachiso Nwadike , Samuele Poppi , Jie Zhang , Karthik Nandakumar , Neil Gong , Nils Lukas

As LLMs become commonplace, machine-generated text has the potential to flood the internet with spam, social media bots, and valueless content. Watermarking is a simple and effective strategy for mitigating such harms by enabling the…

Digital watermarking is a promising solution for mitigating some of the risks arising from the misuse of automatically generated text. These approaches either embed non-specific watermarks to allow for the detection of any text generated by…

密码学与安全 · 计算机科学 2025-06-23 Zihao Fu , Chris Russell

Emerging technologies drive the ongoing transformation of Intelligent Transportation Systems (ITS). This transformation has given rise to cybersecurity concerns, among which data poisoning attack emerges as a new threat as ITS increasingly…

密码学与安全 · 计算机科学 2024-07-24 Feilong Wang , Xin Wang , Xuegang Ban

Detecting whether copyright holders' works were used in LLM pretraining is poised to be an important problem. This work proposes using data watermarks to enable principled detection with only black-box model access, provided that the…

密码学与安全 · 计算机科学 2024-08-20 Johnny Tian-Zheng Wei , Ryan Yixiang Wang , Robin Jia

Watermarking has been widely adopted for protecting the intellectual property (IP) of Deep Neural Networks (DNN) to defend the unauthorized distribution. Unfortunately, the popular data-poisoning DNN watermarking scheme relies on target…

密码学与安全 · 计算机科学 2022-10-18 Run Wang , Jixing Ren , Boheng Li , Tianyi She , Chenhao Lin , Liming Fang , Jing Chen , Chao Shen , Lina Wang