中文
相关论文

相关论文: Watermarking Counterfactual Explanations

200 篇论文

With the rapid rise of generative AI and synthetic media, distinguishing AI-generated images from real ones has become crucial in safeguarding against misinformation and ensuring digital authenticity. Traditional watermarking techniques…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Vinu Sankar Sadasivan , Mehrdad Saberi , Soheil Feizi

Despite progress in watermarking algorithms for large language models (LLMs), real-world deployment remains limited. We argue that this gap stems from misaligned incentives among LLM providers, platforms, and end users, which manifest as…

密码学与安全 · 计算机科学 2026-04-22 Yepeng Liu , Xuandong Zhao , Dawn Song , Gregory W. Wornell , Yuheng Bu

Large language models (LLMs) are pre-trained and post-trained on vast amounts of loosely curated data, raising the possibility that these models may have been trained on proprietary datasets or the same benchmarks used for evaluation. This…

机器学习 · 计算机科学 2026-05-11 Pengrun Huang , Kamalika Chaudhuri , Yu-Xiang Wang

Potential harms of large language models can be mitigated by watermarking model output, i.e., embedding signals into generated text that are invisible to humans but algorithmically detectable from a short span of tokens. We propose a…

机器学习 · 计算机科学 2024-05-03 John Kirchenbauer , Jonas Geiping , Yuxin Wen , Jonathan Katz , Ian Miers , Tom Goldstein

Deepfakes pose significant security and privacy threats through malicious facial manipulations. While robust watermarking can aid in authenticity verification and source tracking, existing methods often lack the sufficient robustness…

计算机视觉与模式识别 · 计算机科学 2025-10-13 Chen Sun , Haiyang Sun , Zhiqing Guo , Yunfeng Diao , Liejun Wang , Dan Ma , Gaobo Yang , Keqin Li

Deepfakes and manipulated media are becoming a prominent threat due to the recent advances in realistic image and video synthesis techniques. There have been several attempts at combating Deepfakes using machine learning classifiers.…

计算机视觉与模式识别 · 计算机科学 2022-04-06 Paarth Neekhara , Shehzeen Hussain , Xinqiao Zhang , Ke Huang , Julian McAuley , Farinaz Koushanfar

The most effective techniques to detect LLM-generated text rely on inserting a detectable signature -- or watermark -- during the model's decoding process. Most existing watermarking methods require access to the underlying LLM's logits,…

机器学习 · 计算机科学 2024-10-14 Yapei Chang , Kalpesh Krishna , Amir Houmansadr , John Wieting , Mohit Iyyer

Watermarking large language models (LLMs) is vital for preventing their misuse, including the fabrication of fake news, plagiarism, and spam. It is especially important to watermark LLM-generated code, as it often contains intellectual…

密码学与安全 · 计算机科学 2025-12-18 Li Lin , Siyuan Xin , Yang Cao , Xiaochun Cao

As large language models (LLMs) generate texts with increasing fluency and realism, there is a growing need to identify the source of texts to prevent the abuse of LLMs. Text watermarking techniques have proven reliable in distinguishing…

计算与语言 · 计算机科学 2024-04-04 Lean Wang , Wenkai Yang , Deli Chen , Hao Zhou , Yankai Lin , Fandong Meng , Jie Zhou , Xu Sun

Counterfactual examples (CFs) are one of the most popular methods for attaching post-hoc explanations to machine learning (ML) models. However, existing CF generation methods either exploit the internals of specific models or depend on each…

机器学习 · 计算机科学 2023-08-10 Ziheng Chen , Fabrizio Silvestri , Jia Wang , He Zhu , Hongshik Ahn , Gabriele Tolomei

LLM watermarking has attracted attention as a promising way to detect AI-generated content, with some works suggesting that current schemes may already be fit for deployment. In this work we dispute this claim, identifying watermark…

机器学习 · 计算机科学 2024-06-25 Nikola Jovanović , Robin Staab , Martin Vechev

Malicious Deepfakes have led to a sharp conflict over distinguishing between genuine and forged faces. Although many countermeasures have been developed to detect Deepfakes ex-post, undoubtedly, passive forensics has not considered any…

计算机视觉与模式识别 · 计算机科学 2023-08-02 Xiaoshuai Wu , Xin Liao , Bo Ou

Recent advancements in large language models (LLMs) have highlighted the risk of misusing them, raising the need for accurate detection of LLM-generated content. In response, a viable solution is to inject imperceptible identifiers into…

计算与语言 · 计算机科学 2025-02-11 Minjia Mao , Dongjun Wei , Zeyu Chen , Xiao Fang , Michael Chau

Watermarking becomes one of the pivotal solutions to trace and verify the origin of synthetic images generated by artificial intelligence models, but it is not free of risks. Recent studies demonstrate the capability to forge watermarks…

多媒体 · 计算机科学 2025-06-09 Chaoyi Zhu , Zaitang Li , Renyi Yang , Robert Birke , Pin-Yu Chen , Tsung-Yi Ho , Lydia Y. Chen

As deep learning (DL) models are widely and effectively used in Machine Learning as a Service (MLaaS) platforms, there is a rapidly growing interest in DL watermarking techniques that can be used to confirm the ownership of a particular…

密码学与安全 · 计算机科学 2024-11-22 Mikhail Pautov , Nikita Bogdanov , Stanislav Pyatkin , Oleg Rogov , Ivan Oseledets

Counterfactual explanations offer an intuitive way to interpret graph neural networks (GNNs) by identifying minimal changes that alter a model's prediction, thereby answering "what must differ for a different outcome?". In this work, we…

机器学习 · 计算机科学 2026-02-09 Yu Zhang , Sean Bin Yang , Arijit Khan , Cuneyt Gurcan Akcora

Watermarking data for source tracking applications by its owner can be unfair for recipients because the data owner may redistribute the same watermarked data to many users. Hence, each data recipient should know the watermark embedded in…

密码学与安全 · 计算机科学 2023-02-06 Mesfer Mohammed Alqarni

Generative models that can produce realistic images have improved significantly in recent years. The quality of the generated content has increased drastically, so sometimes it is very difficult to distinguish between the real images and…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Mikhail Pautov , Danil Ivanov , Andrey V. Galichin , Oleg Rogov , Ivan Oseledets

Watermarking has been widely adopted for protecting the intellectual property (IP) of Deep Neural Networks (DNN) to defend the unauthorized distribution. Unfortunately, the popular data-poisoning DNN watermarking scheme relies on target…

密码学与安全 · 计算机科学 2022-10-18 Run Wang , Jixing Ren , Boheng Li , Tianyi She , Chenhao Lin , Liming Fang , Jing Chen , Chao Shen , Lina Wang

Watermarking is an effective way to trace model-generated content. Current watermark methods cannot resist forgery attacks, such as a deceptive claim that the model-generated content is a response to a fabricated prompt. None of them can be…

密码学与安全 · 计算机科学 2024-12-30 Minhao Bai