中文
相关论文

相关论文: Analyzing and Evaluating Unbiased Language Model W…

200 篇论文

A critical yet frequently overlooked challenge in the field of deepfake detection is the lack of a standardized, unified, comprehensive benchmark. This issue leads to unfair performance comparisons and potentially misleading results.…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Zhiyuan Yan , Yong Zhang , Xinhang Yuan , Siwei Lyu , Baoyuan Wu

Watermarking generative content serves as a vital tool for authentication, ownership protection, and mitigation of potential misuse. Existing watermarking methods face the challenge of balancing robustness and concealment. They empirically…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Huayang Huang , Yu Wu , Qian Wang

Large language models (LLMs) are pre-trained and post-trained on vast amounts of loosely curated data, raising the possibility that these models may have been trained on proprietary datasets or the same benchmarks used for evaluation. This…

机器学习 · 计算机科学 2026-05-11 Pengrun Huang , Kamalika Chaudhuri , Yu-Xiang Wang

This work introduces \textbf{VideoMark}, a distortion-free robust watermarking framework for video diffusion models. As diffusion models excel in generating realistic videos, reliable content attribution is increasingly critical. However,…

密码学与安全 · 计算机科学 2025-11-18 Xuming Hu , Hanqian Li , Jungang Li , Yu Huang , Shuliang Liu , Qi Zheng , Junhao Chen , Aiwei Liu

Existing watermarking algorithms are vulnerable to paraphrase attacks because of their token-level design. To address this issue, we propose SemStamp, a robust sentence-level semantic watermarking algorithm based on locality-sensitive…

In practical application, the widespread deployment of diffusion models often necessitates substantial investment in training. As diffusion models find increasingly diverse applications, concerns about potential misuse highlight the…

计算机视觉与模式识别 · 计算机科学 2024-10-31 Jijia Yang , Sen Peng , Xiaohua Jia

Despite significant progress in designing powerful adversarial evasion attacks for robustness verification, the evaluation of these methods often remains inconsistent and unreliable. Many assessments rely on mismatched models, unverified…

密码学与安全 · 计算机科学 2025-07-08 Antonio Emanuele Cinà , Maura Pintor , Luca Demetrio , Ambra Demontis , Battista Biggio , Fabio Roli

Attack detection and mitigation strategies for cyberphysical systems (CPS) are an active area of research, and researchers have developed a variety of attack-detection tools such as dynamic watermarking. However, such methods often make…

系统与控制 · 电气工程与系统科学 2020-04-01 Matt Olfat , Stephen Sloan , Pedro Hespanhol , Matt Porter , Ram Vasudevan , Anil Aswani

Watermarking is a key technique for detecting AI-generated text. In this work, we study its vulnerabilities and introduce the Smoothing Attack, a novel watermark removal method. By leveraging the relationship between the model's confidence…

机器学习 · 计算机科学 2025-02-06 Hongyan Chang , Hamed Hassani , Reza Shokri

This study investigates the robustness of image classifiers to text-guided corruptions. We utilize diffusion models to edit images to different domains. Unlike other works that use synthetic or hand-picked data for benchmarking, we use…

计算机视觉与模式识别 · 计算机科学 2023-08-01 Mohammadreza Mofayezi , Yasamin Medghalchi

With the growth of editing and sharing images through the internet, the importance of protecting the images' authorship has increased. Robust watermarking is a known approach to maintaining copyright protection. Robustness and…

计算机视觉与模式识别 · 计算机科学 2022-12-27 Arezoo PariZanganeh , Ghazaleh Ghorbanzadeh , Zahra Nabizadeh ShahreBabak , Nader Karimi , Shadrokh Samavi

Text watermarking plays a crucial role in ensuring the traceability and accountability of large language model (LLM) outputs and mitigating misuse. While promising, most existing methods assume perfect pseudorandomness. In practice,…

统计理论 · 数学 2026-01-21 T. Tony Cai , Xiang Li , Qi Long , Weijie J. Su , Garrett G. Wen

Imperceptible digital watermarking is important in copyright protection, misinformation prevention, and responsible generative AI. We propose TrustMark - a GAN-based watermarking method with novel design in architecture and spatio-spectra…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Tu Bui , Shruti Agarwal , John Collomosse

Generated speech achieves human-level naturalness but escalates security risks of misuse. However, existing watermarking methods fail to reconcile fidelity with robustness, as they rely either on simple superposition in the noise space or…

密码学与安全 · 计算机科学 2026-02-02 Weizhi Liu , Yue Li , Zhaoxia Yin

Large language model (LLM) unlearning is critical in real-world applications where it is necessary to efficiently remove the influence of private, copyrighted, or harmful data from some users. Existing utility-centric unlearning metrics…

Gradient Boosting Decision Trees (GBDTs) are widely used in industry and academia for their high accuracy and efficiency, particularly on structured data. However, watermarking GBDT models remains underexplored compared to neural networks.…

人工智能 · 计算机科学 2025-11-14 Jun Woo Chung , Yingjie Lao , Weijie Zhao

In this study, we investigate the vulnerability of image watermarks to diffusion-model-based image editing, a challenge exacerbated by the computational cost of accessing gradient information and the closed-source nature of many diffusion…

计算机视觉与模式识别 · 计算机科学 2024-06-07 Minzhou Pan , Yi Zeng , Xue Lin , Ning Yu , Cho-Jui Hsieh , Peter Henderson , Ruoxi Jia

Ethical concerns surrounding copyright protection and inappropriate content generation pose challenges for the practical implementation of diffusion models. One effective solution involves watermarking the generated images. Existing methods…

计算机视觉与模式识别 · 计算机科学 2025-05-14 Zijin Yang , Xin Zhang , Kejiang Chen , Kai Zeng , Qiyi Yao , Han Fang , Weiming Zhang , Nenghai Yu

Watermarking LLM-generated text is critical for content attribution and misinformation prevention. However, existing methods compromise text quality, require white-box model access and logit manipulation. These limitations exclude API-based…

计算与语言 · 计算机科学 2026-01-13 Zhuohao Yu , Xingru Jiang , Weizheng Gu , Yidong Wang , Qingsong Wen , Shikun Zhang , Wei Ye

Large Language Models (LLMs) are increasingly integrated into diverse industries, posing substantial security risks due to unauthorized replication and misuse. To mitigate these concerns, robust identification mechanisms are widely…

密码学与安全 · 计算机科学 2024-07-25 Xuhong Wang , Haoyu Jiang , Yi Yu , Jingru Yu , Yilun Lin , Ping Yi , Yingchun Wang , Yu Qiao , Li Li , Fei-Yue Wang