English
Related papers

Related papers: Goal-Based Vision-Language Driving

200 papers

Scene understanding and risk-aware attentions are crucial for human drivers to make safe and effective driving decisions. To imitate this cognitive ability in urban autonomous driving while ensuring the transparency and interpretability, we…

Robotics · Computer Science 2025-07-22 Haichao Liu , Haoren Guo , Pei Liu , Benshan Ma , Yuxiang Zhang , Jun Ma , Tong Heng Lee

Personalized driving refers to an autonomous vehicle's ability to adapt its driving behavior or control strategies to match individual users' preferences and driving styles while maintaining safety and comfort standards. However, existing…

Understanding risk in autonomous driving requires not only perception and prediction, but also high-level reasoning about agent behavior and context. Current Vision Language Model (VLM)-based methods primarily ground agents in static images…

Artificial Intelligence · Computer Science 2026-04-21 Yuan Gao , Mattia Piccinini , Roberto Brusnicki , Yuchen Zhang , Johannes Betz

The deployment of Vision-Language Models (VLMs) in safety-critical domains like autonomous driving (AD) is critically hindered by reliability failures, most notably object hallucination. This failure stems from their reliance on ungrounded,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Zhenguo Zhang , Haohan Zheng , Yishen Wang , Le Xu , Tianchen Deng , Xuefeng Chen , Qu Chen , Bo Zhang , Wuxiong Huang

High-fidelity visual reconstruction and novel-view synthesis are essential for realistic closed-loop evaluation in autonomous driving. While 4D Gaussian Splatting (4DGS) offers a promising balance of accuracy and efficiency, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Haibao Yu , Kuntao Xiao , Jiahang Wang , Ruiyang Hao , Yuxin Huang , Guoran Hu , Haifang Qin , Bowen Jing , Yuntian Bo , Ping Luo

Object-goal navigation has traditionally been limited to ground robots with closed-set object vocabularies. Existing multi-agent approaches depend on precomputed probabilistic graphs tied to fixed category sets, precluding generalization to…

Robotics · Computer Science 2026-03-20 MoniJesu James , Amir Atef Habel , Aleksey Fedoseev , Dzmitry Tsetserokou

Out-of-distribution (OOD) scenarios in autonomous driving pose critical challenges, as planners often fail to generalize beyond their training experience, leading to unsafe or unexpected behavior. Vision-Language Models (VLMs) have shown…

Robotics · Computer Science 2025-08-21 Hayeon Oh

Human drivers adeptly navigate complex scenarios by utilizing rich attentional semantics, but the current autonomous systems struggle to replicate this ability, as they often lose critical semantic information when converting 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Pei Liu , Haipeng Liu , Haichao Liu , Xin Liu , Jinxin Ni , Jun Ma

Autonomous driving systems often infer pedestrian yielding behavior from geometric and kinematic cues alone, limiting their ability to reason about visual scene context and age-dependent behavioral variability. This limitation can produce…

Systems and Control · Electrical Eng. & Systems 2026-04-28 Qingwen Pu , Kun Xie , Yuxiang Liu

Detecting anomalous hazards in visual data, particularly in video streams, is a critical challenge in autonomous driving. Existing models often struggle with unpredictable, out-of-label hazards due to their reliance on predefined object…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Shashank Shriram , Srinivasa Perisetla , Aryan Keskar , Harsha Krishnaswamy , Tonko Emil Westerhof Bossen , Andreas Møgelmose , Ross Greer

Recent years have seen remarkable progress in autonomous driving, yet generalization to long-tail and open-world scenarios remains a major bottleneck for large-scale deployment. To address this challenge, some works use LLMs and VLMs for…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Hao Shao , Letian Wang , Yang Zhou , Yuxuan Hu , Zhuofan Zong , Steven L. Waslander , Wei Zhan , Hongsheng Li

Conventional end-to-end autonomous driving methods often rely on explicit global scene representations, which typically consist of 3D object detection, online mapping, and motion prediction. In contrast, human drivers selectively attend to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Ruiqi Song , Xianda Guo , Yanlun Peng , Qinggong Wei , Hangbin Wu , Long Chen

Vision-language models (VLMs) are increasingly being adopted for end-to-end autonomous driving systems due to their exceptional performance in handling long-tail scenarios. However, current VLM-based approaches suffer from two major…

Robotics · Computer Science 2026-03-31 Yuqi Ye , Zijian Zhang , Junhong Lin , Shangkun Sun , Changhao Peng , Wei Gao

Feed-forward surround-view autonomous driving scene reconstruction offers fast, generalizable inference ability, which faces the core challenge of ensuring generalization while elevating novel view quality. Due to the surround-view with…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Junhong Lin , Kangli Wang , Shunzhou Wang , Songlin Fan , Ge Li , Wei Gao

While reasoning technology like Chain of Thought (CoT) has been widely adopted in Vision Language Action (VLA) models, it demonstrates promising capabilities in end to end autonomous driving. However, recent efforts to integrate CoT…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Yuechen Luo , Fang Li , Shaoqing Xu , Zhiyi Lai , Lei Yang , Qimao Chen , Ziang Luo , Zixun Xie , Shengyin Jiang , Jiaxin Liu , Long Chen , Bing Wang , Zhi-xin Yang

Recent advancements in autonomous driving (AD) have explored the use of vision-language models (VLMs) within visual question answering (VQA) frameworks for direct driving decision-making. However, these approaches often depend on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Xin Hu , Taotao Jing , Renran Tian , Zhengming Ding

Research interest in end-to-end autonomous driving has surged owing to its fully differentiable design integrating modular tasks, i.e. perception, prediction and planing, which enables optimization in pursuit of the ultimate goal. Despite…

Artificial Intelligence · Computer Science 2025-06-04 Anqing Jiang , Yu Gao , Zhigang Sun , Yiru Wang , Jijun Wang , Jinghao Chai , Qian Cao , Yuweng Heng , Hao Jiang , Yunda Dong , Zongzheng Zhang , Xianda Guo , Hao Sun , Hao Zhao

Vision-Language-Action (VLA) models are advancing autonomous driving by replacing modular pipelines with unified end-to-end architectures. However, current VLAs face two expensive requirements: (1) massive dataset collection, and (2) dense…

Artificial Intelligence · Computer Science 2026-02-27 Ishaan Rawal , Shubh Gupta , Yihan Hu , Wei Zhan

Autonomous vehicles rely on deep neural networks (DNNs) for traffic sign recognition, lane centering, and vehicle detection, yet these models are vulnerable to attacks that induce misclassification and threaten safety. Existing defenses…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Pedram MohajerAnsari , Amir Salarpour , Michael Kühr , Siyu Huang , Mohammad Hamad , Sebastian Steinhorst , Habeeb Olufowobi , Bing Li , Mert D. Pesé

Autonomous driving presents many challenges due to the large number of scenarios the autonomous vehicle (AV) may encounter. End-to-end deep learning models are comparatively simplistic models that can handle a broad set of scenarios.…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Zhongying CuiZhu , Francois Charette , Amin Ghafourian , Debo Shi , Matthew Cui , Anjali Krishnamachar , Iman Soltani