English
Related papers

Related papers: Waterfall: Framework for Robust and Scalable Text …

200 papers

The vast amounts of digital content captured from the real world or AI-generated media necessitate methods for copyright protection, traceability, or data provenance verification. Digital watermarking serves as a crucial approach to address…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Vitaliy Kinakh , Brian Pulfer , Yury Belousov , Pierre Fernandez , Teddy Furon , Slava Voloshynovskiy

We propose a methodology for planting watermarks in text from an autoregressive language model that are robust to perturbations without changing the distribution over text up to a certain maximum generation budget. We generate watermarked…

Machine Learning · Computer Science 2024-06-07 Rohith Kuditipudi , John Thickstun , Tatsunori Hashimoto , Percy Liang

The capabilities of large language models have grown significantly in recent years and so too have concerns about their misuse. It is important to be able to distinguish machine-generated text from human-authored content. Prior works have…

Cryptography and Security · Computer Science 2024-10-15 Julien Piet , Chawin Sitawarin , Vivian Fang , Norman Mu , David Wagner

Large language models (LLMs) have achieved remarkable success across a wide range of natural language processing tasks, demonstrating human-level performance in text generation, reasoning, and question answering. However, training such…

Cryptography and Security · Computer Science 2025-11-17 Yanbo Dai , Zongjie Li , Zhenlan Ji , Shuai Wang

Large Language Models (LLMs) are increasingly fine-tuned on smaller, domain-specific datasets to improve downstream performance. These datasets often contain proprietary or copyrighted material, raising the need for reliable safeguards…

Computation and Language · Computer Science 2025-10-06 Jingqi Zhang , Ruibo Chen , Yingqing Yang , Peihua Mai , Heng Huang , Yan Pang

Deep Learning (DL) models have caused a paradigm shift in our ability to comprehend raw data in various important fields, ranging from intelligence warfare and healthcare to autonomous transportation and automated manufacturing. A practical…

Cryptography and Security · Computer Science 2018-06-04 Bita Darvish Rouhani , Huili Chen , Farinaz Koushanfar

The rapid advancement of generative AI has underscored the critical need for identifying image ownership and protecting copyrights. This makes post-processing image watermarking an essential tool -- it involves embedding a specific…

Cryptography and Security · Computer Science 2026-05-12 Xinyu Zhang , Ziping Dong , Qingyu Liu , Yuan Hong , Zhongjie Ba , Kui Ren

The widespread use of Large Language Models (LLMs), celebrated for their ability to generate human-like text, has raised concerns about misinformation and ethical implications. Addressing these concerns necessitates the development of…

Computation and Language · Computer Science 2024-03-28 Wissam Antoun , Benoît Sagot , Djamé Seddah

This paper introduces RoSeMary, the first-of-its-kind ML/Crypto codesign watermarking framework that regulates LLM-generated code to avoid intellectual property rights violations and inappropriate misuse in software development.…

Cryptography and Security · Computer Science 2025-02-11 Ruisi Zhang , Neusha Javidnia , Nojan Sheybani , Farinaz Koushanfar

Generation-time text watermarking embeds statistical signals into text for traceability of AI-generated content. We explore *post-hoc watermarking* where an LLM rewrites existing text while applying generation-time watermarking, to protect…

Natural language processing (NLP) technology has shown great commercial value in applications such as sentiment analysis. But NLP models are vulnerable to the threat of pirated redistribution, damaging the economic interests of model…

Cryptography and Security · Computer Science 2022-11-21 Long Dai , Jiarong Mao , Xuefeng Fan , Xiaoyi Zhou

We investigate the radioactivity of text generated by large language models (LLM), i.e. whether it is possible to detect that such synthetic input was used to train a subsequent LLM. Current methods like membership inference or active IP…

Cryptography and Security · Computer Science 2024-10-29 Tom Sander , Pierre Fernandez , Alain Durmus , Matthijs Douze , Teddy Furon

Recently, text watermarking algorithms for large language models (LLMs) have been proposed to mitigate the potential harms of text generated by LLMs, including fake news and copyright issues. However, current watermark detection algorithms…

Computation and Language · Computer Science 2024-05-28 Aiwei Liu , Leyi Pan , Xuming Hu , Shu'ang Li , Lijie Wen , Irwin King , Philip S. Yu

The expansion of the open source community and the rise of large language models have raised ethical and security concerns on the distribution of source code, such as misconduct on copyrighted code, distributions without proper licenses, or…

Cryptography and Security · Computer Science 2024-01-03 Borui Yang , Wei Li , Liyao Xiang , Bo Li

We consider the emerging problem of identifying the presence and use of watermarking schemes in widely used, publicly hosted, closed source large language models (LLMs). We introduce a suite of baseline algorithms for identifying watermarks…

Machine Learning · Computer Science 2023-05-31 Leonard Tang , Gavin Uberti , Tom Shlomi

With the rise of Machine Learning as a Service (MLaaS) platforms,safeguarding the intellectual property of deep learning models is becoming paramount. Among various protective measures, trigger set watermarking has emerged as a flexible and…

Cryptography and Security · Computer Science 2024-04-23 Hongyu Zhu , Sichu Liang , Wentao Hu , Fangqi Li , Ju Jia , Shilin Wang

Generative images have proliferated on Web platforms in social media and online copyright distribution scenarios, and semantic watermarking has increasingly been integrated into diffusion models to support reliable provenance tracking and…

Machine Learning · Computer Science 2026-02-26 Zheng Gao , Xiaoyu Li , Zhicheng Bao , Xiaoyan Feng , Jiaojiao Jiang

Watermarking techniques offer a promising way to identify machine-generated content via embedding covert information into the contents generated from language models. A challenge in the domain lies in preserving the distribution of original…

Cryptography and Security · Computer Science 2024-06-26 Yihan Wu , Zhengmian Hu , Junfeng Guo , Hongyang Zhang , Heng Huang

Digital watermarking is a promising solution for mitigating some of the risks arising from the misuse of automatically generated text. These approaches either embed non-specific watermarks to allow for the detection of any text generated by…

Cryptography and Security · Computer Science 2025-06-23 Zihao Fu , Chris Russell

The increasing use of Large Language Models (LLMs) for generating highly coherent and contextually relevant text introduces new risks, including misuse for unethical purposes such as disinformation or academic dishonesty. To address these…

Computation and Language · Computer Science 2024-10-16 Zhenyu Xu , Kun Zhang , Victor S. Sheng