Technical Report on the CVPR 2026@AdvML Workshop Challenge
Abstract
Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. Built on DriveLM-style multi-view visual question answering, the challenge represents each scene with six synchronized camera images and a structured collection of driving-related question-answer pairs. Participants generate adversarial images and suffix-only textual perturbations that induce model responses to deviate from reference answers while preserving image fidelity and limiting textual cost. The competition comprises two phases, with Phase II adding a hidden black-box model to assess transferability. We describe the task design, submission rules, evaluation protocol, and leaderboard results, and then examine five leading submissions for which technical reports were available. Across these reports, several recurring patterns emerge: image-side attacks are favored by the suffix penalty; scene-level, multi-view optimization is more effective than treating views in isolation; QA types and graph structure provide useful priors for allocating attack budget; feature-space objectives can improve black-box transfer; and typographic content embedded in camera images exposes a persistent vulnerability in driving VLAs. These findings provide a practical reference for future robustness evaluation and defense design in multimodal autonomous-driving systems.
Cite
@article{arxiv.2607.11560,
title = {Technical Report on the CVPR 2026@AdvML Workshop Challenge},
author = {Tianyuan Zhang and Zonglei Jing and Jiangfan Liu and Ligong Zhang and Ke Ma and Chengzhi Sun and Xiaohai Xu and Zhirui Zhang and Qianqian Xu and Qingming Huang and Hanyu Fang and Junhua Liu and Zheng Wang and Xiaoliang Liu and Yuanbo Li and Shuai Gui and Bin Wang and Menghe Zheng and Jing Nie and Hanyang Meng and Zeyang Zhang and Xiang Zhang and Yongxuan Zhu and Rui Ding and Hainan Li and Yongkang Zhang and Zhilei Zhu and Xianglong Kong and Jin Hu and Zonghao Ying and Yisong Xiao and Lei Chen and Haotong Qin and Jiakai Wang and Aishan Liu and Ruikai Li and Julia Karbing and Yinpeng Dong and Zhenfei Yin and Shao Jing and Xia Hu and Jingyi Xu and Juntao Dai and Xinyun Chen and Vishal M. Patel and Xianglong Liu and Dawn Song and Alan Yuille and Philip H. S. Torr and Dacheng Tao},
journal= {arXiv preprint arXiv:2607.11560},
year = {2026}
}