English
Related papers

Related papers: SimKey: A Semantically Aware Key Module for Waterm…

200 papers

Watermarking of language model outputs enables statistical detection of model-generated text, which can mitigate harms and misuses of language models. Existing watermarking strategies operate by altering the decoder of an existing language…

Machine Learning · Computer Science 2024-05-03 Chenchen Gu , Xiang Lisa Li , Percy Liang , Tatsunori Hashimoto

This paper presents a survey and taxonomy of LLM fingerprinting and watermarking for identity, ownership verification, provenance, and generated-content attribution. Large language models (LLMs) require substantial investments in data,…

Cryptography and Security · Computer Science 2026-05-29 Bing Liu , Shunping Wang , Yufan Zhu , Xinyi Yu , Jing Huang , Linkang Du , Hongbin Pei , Wei Luo

Watermarking generative-AI systems, such as LLMs, has gained considerable interest, driven by their enhanced capabilities across a wide range of tasks. Although current approaches have demonstrated that small, context-dependent shifts in…

Computation and Language · Computer Science 2024-03-29 Piotr Molenda , Adian Liusie , Mark J. F. Gales

Large language model (LLM) unlearning is critical in real-world applications where it is necessary to efficiently remove the influence of private, copyrighted, or harmful data from some users. Existing utility-centric unlearning metrics…

In this thesis, we develop algorithms with theoretical guarantees for ensuring reliability and accountability of Machine Learning (ML) systems. As ML systems evolve from predictive models to generative models and autonomous agents, the…

Machine Learning · Computer Science 2026-05-12 Carol Xuan Long

Language model (LM) watermarking techniques inject a statistical signal into LM-generated content by substituting the random sampling process with pseudo-random sampling, using watermark keys as the random seed. Among these statistical…

Cryptography and Security · Computer Science 2024-06-06 Yihan Wu , Ruibo Chen , Zhengmian Hu , Yanshuo Chen , Junfeng Guo , Hongyang Zhang , Heng Huang

The rapid adoption of large language models (LLMs), such as GPT-4 and Claude 3.5, underscores the need to distinguish LLM-generated text from human-written content to mitigate the spread of misinformation and misuse in education. One…

Machine Learning · Statistics 2025-11-11 Xingchi Li , Xiaochi Liu , Guanxun Li

Recently, text watermarking algorithms for large language models (LLMs) have been proposed to mitigate the potential harms of text generated by LLMs, including fake news and copyright issues. However, current watermark detection algorithms…

Computation and Language · Computer Science 2024-05-28 Aiwei Liu , Leyi Pan , Xuming Hu , Shu'ang Li , Lijie Wen , Irwin King , Philip S. Yu

Speech watermarking techniques can proactively mitigate the potential harmful consequences of instant voice cloning techniques. These techniques involve the insertion of signals into speech that are imperceptible to humans but can be…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-19 Shengpeng Ji , Ziyue Jiang , Jialong Zuo , Minghui Fang , Yifu Chen , Tao Jin , Zhou Zhao

Text content created by humans or language models is often stolen or misused by adversaries. Tracing text provenance can help claim the ownership of text content or identify the malicious users who distribute misleading content like…

Cryptography and Security · Computer Science 2021-12-16 Xi Yang , Jie Zhang , Kejiang Chen , Weiming Zhang , Zehua Ma , Feng Wang , Nenghai Yu

Natural language generation (NLG) applications have gained great popularity due to the powerful deep learning techniques and large training corpus. The deployed NLG models may be stolen or used without authorization, while watermarking has…

Multimedia · Computer Science 2021-12-13 Tao Xiang , Chunlong Xie , Shangwei Guo , Jiwei Li , Tianwei Zhang

Code Summarization Model (CSM) has been widely used in code production, such as online and web programming for PHP and Javascript. CSMs are essential tools in code production, enhancing software development efficiency and driving innovation…

Cryptography and Security · Computer Science 2025-02-11 Jiale Zhang , Haoxuan Li , Di Wu , Xiaobing Sun , Qinghua Lu , Guodong Long

Large language models (LLMs) have significantly enhanced the usability of AI-generated code, providing effective assistance to programmers. This advancement also raises ethical and legal concerns, such as academic dishonesty or the…

Cryptography and Security · Computer Science 2025-08-04 Boquan Li , Zirui Fu , Mengdi Zhang , Peixin Zhang , Jun Sun , Xingmei Wang

Large pre-trained language models (PLMs) have proven to be a crucial component of modern natural language processing systems. PLMs typically need to be fine-tuned on task-specific downstream datasets, which makes it hard to claim the…

Computation and Language · Computer Science 2023-02-13 Chenxi Gu , Chengsong Huang , Xiaoqing Zheng , Kai-Wei Chang , Cho-Jui Hsieh

Large language model (LLM) watermarks enable authentication of text provenance, curb misuse of machine-generated text, and promote trust in AI systems. Current watermarks operate by changing the next-token predictions output by an LLM. The…

Cryptography and Security · Computer Science 2025-12-03 Dor Tsur , Carol Xuan Long , Claudio Mayrink Verdun , Hsiang Hsu , Chen-Fu Chen , Haim Permuter , Sajani Vithana , Flavio P. Calmon

Large language models (LLMs) excellently generate human-like text, but also raise concerns about misuse in fake news and academic dishonesty. Decoding-based watermark, particularly the GumbelMax-trick-based watermark(GM watermark), is a…

Computation and Language · Computer Science 2024-05-29 Jiayi Fu , Xuandong Zhao , Ruihan Yang , Yuansen Zhang , Jiangjie Chen , Yanghua Xiao

Large language models (LLMs) have demonstrated powerful capabilities in both text understanding and generation. Companies have begun to offer Embedding as a Service (EaaS) based on these LLMs, which can benefit various natural language…

Computation and Language · Computer Science 2023-06-05 Wenjun Peng , Jingwei Yi , Fangzhao Wu , Shangxi Wu , Bin Zhu , Lingjuan Lyu , Binxing Jiao , Tong Xu , Guangzhong Sun , Xing Xie

Watermarking has emerged as a promising way to detect LLM-generated text, by augmenting LLM generations with later detectable signals. Recent work has proposed multiple families of watermarking schemes, several of which focus on preserving…

Cryptography and Security · Computer Science 2025-02-25 Thibaud Gloaguen , Nikola Jovanović , Robin Staab , Martin Vechev

Watermarking has emerged as a promising solution for tracing and authenticating text generated by large language models (LLMs). A common approach to LLM watermarking is to construct a green/red token list and assign higher or lower…

Cryptography and Security · Computer Science 2025-10-27 Li An , Yujian Liu , Yepeng Liu , Yuheng Bu , Yang Zhang , Shiyu Chang

As policy catches up with the capabilities of generative AI, watermarking is central to content provenance efforts. Inference-time watermarks for autoregressive models are unfit for continuous modalities due to discretization…

Machine Learning · Computer Science 2026-05-26 Georgios Milis , Yubin Qin , Yihan Wu , Heng Huang