English
Related papers

Related papers: Tex3D: Objects as Attack Surfaces via Adversarial …

200 papers

Recent advances in large vision-language models (LVLMs) have showcased their remarkable capabilities across a wide range of multimodal vision-language tasks. However, these models remain vulnerable to visual adversarial attacks, which can…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Yudong Zhang , Ruobing Xie , Yiqing Huang , Jiansheng Chen , Xingwu Sun , Zhanhui Kang , Di Wang , Yu Wang

While neural machine translation (NMT) models achieve success in our daily lives, they show vulnerability to adversarial attacks. Despite being harmful, these attacks also offer benefits for interpreting and enhancing NMT models, thus…

Computation and Language · Computer Science 2024-09-10 Yanni Xue , Haojie Hao , Jiakai Wang , Qiang Sheng , Renshuai Tao , Yu Liang , Pu Feng , Xianglong Liu

Multimodal large language models (MLLMs) have advanced the capabilities to interpret and act on visual input in 3D environments, empowering diverse applications such as robotics and situated conversational agents. When MLLMs reason over…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Zhuoheng Li , Ying Chen

Vision-language-action models (VLAs) show potential as generalist robot policies. However, these models pose extreme safety challenges during real-world deployment, including the risk of harm to the environment, the robot itself, and…

Robotics · Computer Science 2026-04-21 Borong Zhang , Yuhao Zhang , Jiaming Ji , Yingshan Lei , Yishuai Cai , Josef Dai , Yuanpei Chen , Yaodong Yang

On-device Vision-Language Models (VLMs) promise data privacy via local execution. However, we show that the architectural shift toward Dynamic High-Resolution preprocessing (e.g., AnyRes) introduces an inherent algorithmic side-channel.…

Cryptography and Security · Computer Science 2026-03-30 Eyal Hadad , Mordechai Guri

Physical adversarial attack methods expose the vulnerabilities of deep neural networks and pose a significant threat to safety-critical scenarios such as autonomous driving. Camouflage-based physical attack is a more promising approach…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Tianrui Lou , Xiaojun Jia , Siyuan Liang , Jiawei Liang , Ming Zhang , Yanjun Xiao , Xiaochun Cao

The rapid evolution of Embodied AI has enabled Vision-Language-Action (VLA) models to excel in multimodal perception and task execution. However, applying Reinforcement Learning (RL) to these massive models in large-scale distributed…

Artificial Intelligence · Computer Science 2026-05-15 Yucheng Guo , Yongjian Guo , Zhong Guan , Wen Huang , Haoran Sun , Haodong Yue , Xiaolong Xiang , Shuai Di , Zhen Sun , Luqiao Wang , Junwu Xiong , Yicheng Gong

The rapid progress of auto-regressive vision-language models (VLMs) has inspired growing interest in vision-language-action models (VLA) for robotic manipulation. Recently, masked diffusion models, a paradigm distinct from autoregressive…

Robotics · Computer Science 2025-09-11 Yuqing Wen , Hebei Li , Kefan Gu , Yucheng Zhao , Tiancai Wang , Xiaoyan Sun

Over the past decade, deep learning has revolutionized conventional tasks that rely on hand-craft feature extraction with its strong feature learning capability, leading to substantial enhancements in traditional tasks. However, deep neural…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Donghua Wang , Wen Yao , Tingsong Jiang , Guijian Tang , Xiaoqian Chen

Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving robot's generalization and robustness. OpenAI's recent…

Neural networks are vulnerable to adversarial attacks -- small visually imperceptible crafted noise which when added to the input drastically changes the output. The most effective method of defending against these adversarial attacks is to…

Many robotic manipulation tasks require sensing and responding to force signals such as torque to assess whether the task has been successfully completed and to enable closed-loop control. However, current Vision-Language-Action (VLA)…

Robotics · Computer Science 2025-09-10 Zongzheng Zhang , Haobo Xu , Zhuo Yang , Chenghao Yue , Zehao Lin , Huan-ang Gao , Ziwei Wang , Hao Zhao

Vision-Language-Action (VLA) models map multimodal perception and language instructions to executable robot actions, making them particularly vulnerable to behavioral backdoor manipulation: a hidden trigger introduced during training can…

Cryptography and Security · Computer Science 2026-03-10 Zonghuan Xu , Jiayu Li , Yunhan Zhao , Xiang Zheng , Xingjun Ma , Yu-Gang Jiang

Vision-Language-Action (VLA) models are vulnerable to adversarial attacks, yet universal and transferable attacks remain underexplored, as most existing patches overfit to a single model and fail in black-box settings. To address this gap,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Hui Lu , Yi Yu , Yiming Yang , Chenyu Yi , Qixin Zhang , Bingquan Shen , Alex C. Kot , Xudong Jiang

Large-scale Video Foundation Models (VFMs) has significantly advanced various video-related tasks, either through task-specific models or Multi-modal Large Language Models (MLLMs). However, the open accessibility of VFMs also introduces…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Hui Lu , Yi Yu , Song Xia , Yiming Yang , Deepu Rajan , Boon Poh Ng , Alex Kot , Xudong Jiang

The growing success of Vision-Language-Action (VLA) models stems from the promise that pretrained Vision-Language Models (VLMs) can endow agents with transferable world knowledge and vision-language (VL) grounding, laying a foundation for…

Machine Learning · Computer Science 2025-10-30 Nikita Kachaev , Mikhail Kolosov , Daniil Zelezetsky , Alexey K. Kovalev , Aleksandr I. Panov

We present Adversarial Object Fusion (AdvOF), a novel attack framework targeting vision-and-language navigation (VLN) agents in service-oriented environments by generating adversarial 3D objects. While foundational models like Large…

Cryptography and Security · Computer Science 2025-05-30 Chunlong Xie , Jialing He , Shangwei Guo , Jiacheng Wang , Shudong Zhang , Tianwei Zhang , Tao Xiang

Vision-Language-Action (VLA) models have become foundational to modern embodied AI systems. By integrating visual perception, language understanding, and action planning, they enable general-purpose task execution across diverse…

Robotics · Computer Science 2026-02-03 Jianyi Zhou , Yujie Wei , Ruichen Zhen , Bo Zhao , Xiaobo Xia , Rui Shao , Xiu Su , Shuo Yang

Vision-language-action (VLA) models perform well on training-seen robotic tasks but struggle to generalize to unseen scenes and objects. A key limitation lies in their implicit visual representations, which entangle object appearance,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Hanyu Zhou , Chuanhao Ma , Gim Hee Lee

Vision-language pre-training (VLP) models are vulnerable to adversarial examples, particularly in black-box scenarios. Existing multimodal attacks often suffer from limited perturbation diversity and unstable multi-stage pipelines. To…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Wutao Chen , Huaqin Zou , Chen Wan , Lifeng Huang
‹ Prev 1 4 5 6 7 8 10 Next ›