English
Related papers

Related papers: MASH: Evading Black-Box AI-Generated Text Detector…

200 papers

Black box attacks, where adversaries have limited knowledge of the target model, pose a significant threat to machine learning systems. Adversarial examples generated with a substitute model often suffer from limited transferability to the…

Machine Learning · Computer Science 2024-10-22 Bar Avraham , Yisroel Mirsky

Large language models (LLMs) have shown the ability to produce fluent and cogent content, presenting both productivity opportunities and societal risks. To build trustworthy AI systems, it is imperative to distinguish between…

Computation and Language · Computer Science 2024-12-17 Guangsheng Bao , Yanbin Zhao , Zhiyang Teng , Linyi Yang , Yue Zhang

Machine-generated text (MGT) detection is critical for regulating online information ecosystems, yet existing detectors often underperform in few-shot settings and remain vulnerable to adversarial, humanizing attacks. To build accurate and…

Cryptography and Security · Computer Science 2026-05-05 Wenjing Duan , Qi Zhou , Yuanfan Li

Watermark has been widely deployed by industry to detect AI-generated images. The robustness of such watermark-based detector against evasion attacks in the white-box and black-box settings is well understood in the literature. However, the…

Cryptography and Security · Computer Science 2025-02-20 Yuepeng Hu , Zhengyuan Jiang , Moyang Guo , Neil Zhenqiang Gong

The prosperous development of Artificial Intelligence-Generated Content (AIGC) has brought people's anxiety about the spread of false information on social media. Designing detectors for filtering is an effective defense method, but most…

Cryptography and Security · Computer Science 2025-12-11 Xiaojing Chen , Dan Li , Lijun Peng , Jun YanŁetter , Zhiqing Guo , Junyang Chen , Xiao Lan , Zhongjie Ba , Yunfeng DiaoŁetter

Recently, text-to-image diffusion models have been widely used for style mimicry and personalized customization through methods such as DreamBooth and Textual Inversion. This has raised concerns about intellectual property protection and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Yanjie Li , Wenxuan Zhang , Xinqi Lyu , Yihao Liu , Bin Xiao

Current multi-task adversarial text attacks rely on abundant access to shared internal features and numerous queries, often limited to a single task type. As a result, these attacks are less effective against practical scenarios involving…

Cryptography and Security · Computer Science 2025-08-15 Wenqiang Wang , Yan Xiao , Hao Lin , Yangshijie Zhang , Xiaochun Cao

We find that large language models (LLMs) are more likely to modify human-written text than AI-generated text when tasked with rewriting. This tendency arises because LLMs often perceive AI-generated text as high-quality, leading to fewer…

Computation and Language · Computer Science 2024-04-16 Chengzhi Mao , Carl Vondrick , Hao Wang , Junfeng Yang

Modern commercial antivirus systems increasingly rely on machine learning to keep up with the rampant inflation of new malware. However, it is well-known that machine learning models are vulnerable to adversarial examples (AEs). Previous…

Cryptography and Security · Computer Science 2021-05-03 Wei Song , Xuezixiang Li , Sadia Afroz , Deepali Garg , Dmitry Kuznetsov , Heng Yin

In recent years, there have been significant advancements in the development of Large Language Models (LLMs). While their practical applications are now widespread, their potential for misuse, such as generating fake news and committing…

Artificial Intelligence · Computer Science 2024-04-01 Kaito Taguchi , Yujie Gu , Kouichi Sakurai

Smishing, or SMS-based phishing, poses an increasing threat to mobile users by mimicking legitimate communications through culturally adapted, concise, and deceptive messages, which can result in the loss of sensitive data or financial…

Machine Learning · Computer Science 2025-06-04 Shaghayegh Hosseinpour , Sanchari Das

Detecting vehicles in aerial images is difficult due to complex backgrounds, small object sizes, shadows, and occlusions. Although recent deep learning advancements have improved object detection, these models remain susceptible to…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Mikael Yeghiazaryan , Sai Abhishek Siddhartha Namburu , Emily Kim , Stanislav Panev , Celso de Melo , Fernando De la Torre , Jessica K. Hodgins

Deep neural network-based classifiers are prone to errors when processing adversarial examples (AEs). AEs are minimally perturbed input data undetectable to humans posing significant risks to security-dependent applications. Hence,…

Cryptography and Security · Computer Science 2026-01-05 Fumiya Morimoto , Ryuto Morita , Satoshi Ono

Retrieval-Augmented Generation (RAG) systems enhance response credibility and traceability by displaying reference contexts, but this transparency simultaneously introduces a novel black-box attack vector. Existing document poisoning…

Computation and Language · Computer Science 2026-01-27 Runqi Sui

It is well known that fraudulent reviews cast doubt on the legitimacy and dependability of online purchases. The most recent development that leads customers towards darkness is the appearance of human reviews in computer-generated (CG)…

Computation and Language · Computer Science 2025-12-01 Shabbir Anees , Anshuman , Ayush Chaurasia , Prathmesh Bogar

Large language models (LLMs) have advanced to a point that even humans have difficulty discerning whether a text was generated by another human, or by a computer. However, knowing whether a text was produced by human or artificial…

Computation and Language · Computer Science 2025-04-15 Kathleen C. Fraser , Hillary Dawkins , Svetlana Kiritchenko

Detecting machine-generated text (MGT) from contemporary Large Language Models (LLMs) is increasingly crucial amid risks like disinformation and threats to academic integrity. Existing zero-shot detection paradigms, despite their…

Computation and Language · Computer Science 2025-08-19 Yue Wang , Liesheng Wei , Yuxiang Wang

The increasing capability of large language models (LLMs) to generate fluent long-form texts is presenting new challenges in distinguishing machine-generated outputs from human-written ones, which is crucial for ensuring authenticity and…

Computation and Language · Computer Science 2024-10-08 Yufei Tian , Zeyu Pan , Nanyun Peng

Autonomous AI agents executing multi-step tool sequences face semantic attacks that manifest in behavioral traces rather than isolated prompts. A critical challenge is cross-attack generalization: can detectors trained on known attack…

Cryptography and Security · Computer Science 2026-01-06 Vignesh Iyer

The efficacy of detectors for texts generated by large language models (LLMs) substantially depends on the availability of large-scale training data. However, white-box zero-shot detectors, which require no such data, are limited by the…

Computation and Language · Computer Science 2025-03-04 Junchao Wu , Runzhe Zhan , Derek F. Wong , Shu Yang , Xuebo Liu , Lidia S. Chao , Min Zhang