English
Related papers

Related papers: MIRAGE: Context-Aware Prompt Injection against Mob…

200 papers

Multimodal AI systems have achieved remarkable performance across a broad range of real-world tasks, yet the mechanisms underlying visual-language reasoning remain surprisingly poorly understood. We report three findings that challenge…

Artificial Intelligence · Computer Science 2026-04-03 Mohammad Asadi , Jack W. O'Sullivan , Fang Cao , Tahoura Nedaee , Kamyar Rajabalifardi , Fei-Fei Li , Ehsan Adeli , Euan Ashley

Web-use agents are rapidly being deployed to automate complex web tasks with extensive browser capabilities. However, these capabilities create a critical and previously unexplored attack surface. This paper demonstrates how attackers can…

Cryptography and Security · Computer Science 2025-10-22 Avishag Shapira , Parth Atulbhai Gandhi , Edan Habler , Asaf Shabtai

We introduce MIRAGE, a new benchmark for multimodal expert-level reasoning and decision-making in consultative interaction settings. Designed for the agriculture domain, MIRAGE captures the full complexity of expert consultations by…

Machine Learning · Computer Science 2026-01-07 Vardhan Dongre , Chi Gui , Shubham Garg , Hooshang Nayyeri , Gokhan Tur , Dilek Hakkani-Tür , Vikram S. Adve

The spreading of AI-generated images (AIGI), driven by advances in generative AI, poses a significant threat to information security and public trust. Existing AIGI detectors, while effective against images in clean laboratory settings,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Cheng Xia , Manxi Lin , Jiexiang Tan , Xiaoxiong Du , Yang Qiu , Junjun Zheng , Xiangheng Kong , Yuning Jiang , Bo Zheng

Web agents powered by vision-language models (VLMs) enable autonomous interaction with web environments by perceiving and acting on both visual and textual webpage content to accomplish user-specified tasks. However, they are highly…

Cryptography and Security · Computer Science 2026-04-15 Yulin Chen , Tri Cao , Haoran Li , Yue Liu , Yibo Li , Yufei He , Le Minh Khoi , Yangqiu Song , Shuicheng Yan , Bryan Hooi

Graphical user interface (GUI) agents powered by multimodal large language models (MLLMs) have shown greater promise for human-interaction. However, due to the high fine-tuning cost, users often rely on open-source GUI agents or APIs…

Computation and Language · Computer Science 2025-05-26 Pengzhou Cheng , Haowen Hu , Zheng Wu , Zongru Wu , Tianjie Ju , Zhuosheng Zhang , Gongshen Liu

Recent advances in vision-language models (VLMs) have significantly enhanced the visual grounding task, which involves locating objects in an image based on natural language queries. Despite these advancements, the security of VLM-based…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Junxian Li , Beining Xu , Simin Chen , Jiatong Li , Jingdi Lei , Haodong Zhao , Di Zhang

Large foundation models are integrated into Computer Use Agents (CUAs), enabling autonomous interaction with operating systems through graphical user interfaces (GUIs) to perform complex tasks. This autonomy introduces serious security…

Artificial Intelligence · Computer Science 2026-01-21 Wenqi Zhang , Yulin Shen , Changyue Jiang , Jiarun Dai , Geng Hong , Xudong Pan

The virtual content in augmented reality (AR) can introduce misleading or harmful information, leading to semantic misunderstandings or user errors. In this work, we focus on visual information manipulation (VIM) attacks in AR, where…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yanming Xiu , Maria Gorlatova

Detecting illicit visual content demands more than image-level NSFW flags; moderators must also know what objects make an image illegal and where those objects occur. We introduce a zero-shot pipeline that simultaneously (i) detects if an…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Sheng Hang , Chaoxiang He , Hongsheng Hu , Hanqing Hu , Bin Benjamin Zhu , Shi-Feng Sun , Dawu Gu , Shuo Wang

Computer-Use Agents (CUAs) with full system access enable powerful task automation but pose significant security and privacy risks due to their ability to manipulate files, access user data, and execute arbitrary commands. While prior work…

Artificial Intelligence · Computer Science 2026-03-03 Tri Cao , Bennett Lim , Yue Liu , Yuan Sui , Yuexin Li , Shumin Deng , Lin Lu , Nay Oo , Shuicheng Yan , Bryan Hooi

Large reasoning models (LRMs) have shown significant progress in test-time scaling through chain-of-thought prompting. Current approaches like search-o1 integrate retrieval augmented generation (RAG) into multi-step reasoning processes but…

Computation and Language · Computer Science 2026-01-21 Kaiwen Wei , Rui Shan , Dongsheng Zou , Jianzhong Yang , Bi Zhao , Junnan Zhu , Jiang Zhong

Augmented reality (AR) enhances user interaction with the real world but also presents vulnerabilities, particularly through Visual Information Manipulation (VIM) attacks. These attacks alter important real-world visual cues, leading to…

Human-Computer Interaction · Computer Science 2025-09-04 Yanming Xiu , Maria Gorlatova

Existing red-teaming studies on GUI agents have important limitations. Adversarial perturbations typically require white-box access, which is unavailable for commercial systems, while prompt injection is increasingly mitigated by stronger…

Cryptography and Security · Computer Science 2026-04-10 Wenkui Yang , Chao Jin , Haisu Zhu , Weilin Luo , Derek Yuen , Kun Shao , Huaibo Huang , Junxian Duan , Jie Cao , Ran He

We introduce MERGE, a system for situational grounding of actors, objects, and events in dynamic human-robot group interactions. Effective collaboration in such settings requires consistent situational awareness, built on persistent…

Mobile agents powered by vision-language models (VLMs) are increasingly adopted for tasks such as UI automation and camera-based assistance. These agents are typically fine-tuned using small-scale, user-collected data, making them…

Cryptography and Security · Computer Science 2025-09-08 Xuan Wang , Siyuan Liang , Zhe Liu , Yi Yu , Aishan Liu , Yuliang Lu , Xitong Gao , Ee-Chien Chang

Vision-language models (VLMs) have revolutionized multimodal AI applications but introduce novel security vulnerabilities that remain largely unexplored. We present the first comprehensive study of steganographic prompt injection attacks…

Cryptography and Security · Computer Science 2025-07-31 Chetan Pathade

Graphical user interface (GUI) agents built on multimodal large language models (MLLMs) have recently demonstrated strong decision-making abilities in screen-based interaction tasks. However, they remain highly vulnerable to pop-up-based…

Cryptography and Security · Computer Science 2026-04-08 Zihe Yan , Jiaping Gui , Zhuosheng Zhang , Gongshen Liu

We present MUG, a novel interactive task for multimodal grounding where a user and an agent work collaboratively on an interface screen. Prior works modeled multimodal UI grounding in one round: the user gives a command and the agent…

Computation and Language · Computer Science 2022-10-03 Tao Li , Gang Li , Jingjie Zheng , Purple Wang , Yang Li

Fusing visual understanding into language generation, Multi-modal Large Language Models (MLLMs) are revolutionizing visual-language applications. Yet, these models are often plagued by the hallucination problem, which involves generating…

Machine Learning · Computer Science 2025-01-28 Yining Wang , Mi Zhang , Junjie Sun , Chenyue Wang , Min Yang , Hui Xue , Jialing Tao , Ranjie Duan , Jiexi Liu