中文
相关论文

相关论文: Towards Optimal Statistical Watermarking

200 篇论文

Large language models (LLMs) can be trained or fine-tuned on data obtained without the owner's consent. Verifying whether a specific LLM was trained on particular data instances or an entire dataset is extremely challenging. Dataset…

计算与语言 · 计算机科学 2025-10-07 Eyal German , Sagiv Antebi , Edan Habler , Asaf Shabtai , Yuval Elovici

Existing watermarking algorithms are vulnerable to paraphrase attacks because of their token-level design. To address this issue, we propose SemStamp, a robust sentence-level semantic watermarking algorithm based on locality-sensitive…

Watermarking has emerged as a promising solution for tracing and authenticating text generated by large language models (LLMs). A common approach to LLM watermarking is to construct a green/red token list and assign higher or lower…

密码学与安全 · 计算机科学 2025-10-27 Li An , Yujian Liu , Yepeng Liu , Yuheng Bu , Yang Zhang , Shiyu Chang

Large language models now draft news, legal analyses, and software code with human-level fluency. At the same time, regulations such as the EU AI Act mandate that each synthetic passage carry an imperceptible, machine-verifiable mark for…

人工智能 · 计算机科学 2025-11-14 Shinwoo Park , Hyejin Park , Hyeseon Ahn , Yo-Sub Han

We review approaches to statistical inference based on randomization. Permutation tests are treated as an important special case. Under a certain group invariance property, referred to as the ``randomization hypothesis,'' randomization…

计量经济学 · 经济学 2025-02-05 David M. Ritzwoller , Joseph P. Romano , Azeem M. Shaikh

The classical binary hypothesis testing problem is revisited. We notice that when one of the hypotheses is composite, there is an inherent difficulty in defining an optimality criterion that is both informative and well-justified. For…

统计理论 · 数学 2021-03-29 Michael Bell , Yuval Kochman

We study a distributed hypothesis testing setup where peripheral nodes send quantized data to the fusion center in a memoryless fashion. The \emph{expected} number of bits sent by each node under the null hypothesis is kept limited. We…

信息论 · 计算机科学 2022-06-27 Yunus Inan , Mert Kayaalp , Ali H. Sayed , Emre Telatar

A new approach to linguistic watermarking of language models is presented in which information is imperceptibly inserted into the output text while preserving its readability and original meaning. A cross-attention mechanism is used to…

计算与语言 · 计算机科学 2024-04-10 Folco Bertini Baldassini , Huy H. Nguyen , Ching-Chung Chang , Isao Echizen

The widespread deployment of high-fidelity generative models has intensified the need for reliable mechanisms for provenance and content authentication. In-processing watermarking, embedding a signature into the generative model's synthesis…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Anirudh Nakra , Min Wu

Typical behavior of the linear programming (LP) problem is studied as a relaxation of the minimum vertex cover, a type of integer programming (IP) problem. A lattice-gas model on the Erd\"os-R\'enyi random graphs of $\alpha$-uniform…

无序系统与神经网络 · 物理学 2016-06-01 Satoshi Takabe , Koji Hukushima

We propose a general method for constructing robust permutation tests under data corruption. The proposed tests effectively control the non-asymptotic type I error under data corruption, and we prove their consistency in power under minimal…

机器学习 · 统计学 2025-04-28 Antonin Schrab , Ilmun Kim

The intellectual property (IP) of Deep neural networks (DNNs) can be easily ``stolen'' by surrogate model attack. There has been significant progress in solutions to protect the IP of DNN models in classification tasks. However, little…

密码学与安全 · 计算机科学 2021-08-06 Jie Zhang , Dongdong Chen , Jing Liao , Han Fang , Zehua Ma , Weiming Zhang , Gang Hua , Nenghai Yu

Active techniques have been introduced to give better detectability performance for cyber-attack diagnosis in cyber-physical systems (CPS). In this paper, switching multiplicative watermarking is considered, whereby we propose an optimal…

密码学与安全 · 计算机科学 2025-12-03 Alexander J. Gallo , Sribalaji C. Anand , André M. H. Teixeira , Riccardo M. G. Ferrari

The rapid growth of Large Language Models (LLMs) raises concerns about distinguishing AI-generated text from human content. Existing watermarking techniques, like \kgw, struggle with low watermark strength and stringent false-positive…

机器学习 · 计算机科学 2025-05-22 Zhuang Li , Qiuping Yi , Zongcheng Ji , Yijian Lu , Yanqi Li , Keyang Xiao , Hongliang Liang

In light of recent advancements in generative AI models, it has become essential to distinguish genuine content from AI-generated one to prevent the malicious usage of fake materials as authentic ones and vice versa. Various techniques have…

计算机视觉与模式识别 · 计算机科学 2024-02-15 Mehrdad Saberi , Vinu Sankar Sadasivan , Keivan Rezaei , Aounon Kumar , Atoosa Chegini , Wenxiao Wang , Soheil Feizi

We present the first undetectable watermarking scheme for generative image models. Undetectability ensures that no efficient adversary can distinguish between watermarked and un-watermarked images, even after making many adaptive queries.…

密码学与安全 · 计算机科学 2025-04-23 Sam Gunn , Xuandong Zhao , Dawn Song

In practical application, the widespread deployment of diffusion models often necessitates substantial investment in training. As diffusion models find increasingly diverse applications, concerns about potential misuse highlight the…

计算机视觉与模式识别 · 计算机科学 2024-10-31 Jijia Yang , Sen Peng , Xiaohua Jia

The sequential multiple testing problem is considered under two generalized error metrics. Under the first one, the probability of at least $k$ mistakes, of any kind, is controlled. Under the second, the probabilities of at least $k_1$…

统计理论 · 数学 2019-02-18 Yanglei Song , Georgios Fellouris

Watermarking of large language models (LLMs) generation embeds an imperceptible statistical pattern within texts, making it algorithmically detectable. Watermarking is a promising method for addressing potential harm and biases from LLMs,…

密码学与安全 · 计算机科学 2024-12-09 Lingjie Chen , Ruizhong Qiu , Siyu Yuan , Zhining Liu , Tianxin Wei , Hyunsik Yoo , Zhichen Zeng , Deqing Yang , Hanghang Tong

A central goal in designing clinical trials is to find the test that maximizes power (or equivalently minimizes required sample size) for finding a false null hypothesis subject to the constraint of type I error. When there is more than one…

统计方法学 · 统计学 2022-09-21 Ruth Heller , Abba Krieger , Saharon Rosset