中文
相关论文

相关论文: Embarrassingly Simple Text Watermarks

200 篇论文

The proliferation of large language models (LLMs) in generating content raises concerns about text copyright. Watermarking methods, particularly logit-based approaches, embed imperceptible identifiers into text to address these challenges.…

计算与语言 · 计算机科学 2025-02-06 Yiyang Luo , Ke Lin , Chao Gu , Jiahui Hou , Lijie Wen , Ping Luo

Watermarking algorithms for Large Language Models (LLMs) effectively identify machine-generated content by embedding and detecting hidden statistical features in text. However, such embedding leads to a decline in text quality, especially…

密码学与安全 · 计算机科学 2025-10-06 Yu Zhang , Shuliang Liu , Xu Yang , Xuming Hu

As large language models (LLMs) become increasingly commonplace, concern about distinguishing between human and AI text increases as well. The growing power of these models is of particular concern to teachers, who may worry that students…

人工智能 · 计算机科学 2024-04-18 James Weichert , Chinecherem Dimobi

Large Language Models (LLMs) have demonstrated exceptional capabilities in natural language understanding and generation. Based on these LLMs, businesses have started to provide Embeddings-as-a-Service (EaaS), offering feature extraction…

计算与语言 · 计算机科学 2025-12-04 Anudeex Shetty

Motivated by the problem of detecting AI-generated text, we consider the problem of watermarking the output of language models with provable guarantees. We aim for watermarks which satisfy: (a) undetectability, a cryptographic notion…

密码学与安全 · 计算机科学 2025-05-29 Noah Golowich , Ankur Moitra

Large Language Models (LLMs) have transformed natural language processing, demonstrating impressive capabilities across diverse tasks. However, deploying these models introduces critical risks related to intellectual property violations and…

密码学与安全 · 计算机科学 2025-12-24 Kieu Dang , Phung Lai , NhatHai Phan , Yelong Shen , Ruoming Jin , Abdallah Khreishah , My T. Thai

Copyright protection and authentication of digital contents has become a significant issue in the current digital epoch with efficient communication mediums such as internet. Plain text is the rampantly used medium used over the internet…

密码学与安全 · 计算机科学 2010-03-10 Zunera Jalil , Anwar M. Mirza , Maria Sabir

Watermarking has become a key technique for proprietary language models, enabling the distinction between AI-generated and human-written text. However, in many real-world scenarios, LLM-generated content may undergo post-generation edits,…

机器学习 · 计算机科学 2025-10-03 Liyan Xie , Muhammad Siddeek , Mohamed Seif , Andrea J. Goldsmith , Mengdi Wang

The rapid growth of Large Language Models (LLMs) has highlighted the pressing need for reliable mechanisms to verify content ownership and ensure traceability. Watermarking offers a promising path forward, but it remains limited by privacy…

密码学与安全 · 计算机科学 2026-01-21 Thomas Fargues , Ye Dong , Tianwei Zhang , Jin-Song Dong

In the rapidly evolving domain of artificial intelligence, safeguarding the intellectual property of Large Language Models (LLMs) is increasingly crucial. Current watermarking techniques against model extraction attacks, which rely on…

密码学与安全 · 计算机科学 2024-05-03 Minhao Bai , Kaiyi Pang , Yongfeng Huang

Natural language generation (NLG) applications have gained great popularity due to the powerful deep learning techniques and large training corpus. The deployed NLG models may be stolen or used without authorization, while watermarking has…

多媒体 · 计算机科学 2021-12-13 Tao Xiang , Chunlong Xie , Shangwei Guo , Jiwei Li , Tianwei Zhang

Large language models (LLMs) are increasingly integrated into real-world personalized applications through retrieval-augmented generation (RAG) mechanisms to supplement their responses with domain-specific knowledge. However, the valuable…

密码学与安全 · 计算机科学 2025-05-26 Junfeng Guo , Yiming Li , Ruibo Chen , Yihan Wu , Chenxi Liu , Yanshuo Chen , Heng Huang

Text watermarking provides an effective solution for identifying synthetic text generated by large language models. However, existing techniques often focus on satisfying specific criteria while ignoring other key aspects, lacking a unified…

密码学与安全 · 计算机科学 2025-03-28 Shuhao Zhang , Bo Cheng , Jiale Han , Yuli Chen , Zhixuan Wu , Changbao Li , Pingli Gu

While watermarking serves as a critical mechanism for LLM provenance, existing secret-key schemes tightly couple detection with injection, requiring access to keys or provider-side scheme-specific detectors for verification. This dependency…

密码学与安全 · 计算机科学 2026-04-14 Zhuoshang Wang , Yubing Ren , Yanan Cao , Fang Fang , Xiaoxue Li , Li Guo

Watermarking (WM) is a critical mechanism for detecting and attributing AI-generated content. Current WM methods for Large Language Models (LLMs) are predominantly tailored for autoregressive (AR) models: They rely on tokens being generated…

计算与语言 · 计算机科学 2026-01-21 Ofek Raban , Ethan Fetaya , Gal Chechik

Watermarking has recently emerged as an effective strategy for detecting the generations of large language models (LLMs). The strength of a watermark typically depends strongly on the entropy afforded by the language model and the set of…

计算与语言 · 计算机科学 2026-02-05 Dara Bahri , John Wieting

To mitigate potential risks associated with language models, recent AI detection research proposes incorporating watermarks into machine-generated text through random vocabulary restrictions and utilizing this information for detection.…

计算与语言 · 计算机科学 2024-02-14 Yu Fu , Deyi Xiong , Yue Dong

Existing watermarking algorithms are vulnerable to paraphrase attacks because of their token-level design. To address this issue, we propose SemStamp, a robust sentence-level semantic watermarking algorithm based on locality-sensitive…

Text watermarking aims to subtly embed statistical signals into text by controlling the Large Language Model (LLM)'s sampling process, enabling watermark detectors to verify that the output was generated by the specified model. The…

机器学习 · 计算机科学 2025-05-13 Yixin Cheng , Hongcheng Guo , Yangming Li , Leonid Sigal

Watermarking has emerged as a crucial technique for detecting and attributing content generated by large language models. While recent advancements have utilized watermark ensembles to enhance robustness, prevailing methods typically…

密码学与安全 · 计算机科学 2026-02-13 Ruibo Chen , Yihan Wu , Xuehao Cui , Jingqi Zhang , Heng Huang