English
Related papers

Related papers: FutureSightDrive: Thinking Visually with Spatio-Te…

200 papers

End-to-end autonomous driving via Vision-Language-Action (VLA) models demands a precarious balance between high-fidelity trajectory planning and efficient inference. Existing paradigms typically fall short: autoregressive (AR) VLAs are…

Computation and Language · Computer Science 2026-05-26 Kewei Zhang , Jin Wang , Sensen Gao , Chengyue Wu , Yulong Cao , Songyang Han , Boris Ivanovic , Langechuan Liu , Marco Pavone , Song Han , Daquan Zhou , Enze Xie

Vision Language Models (VLMs) bridge visual perception and linguistic reasoning. In Autonomous Driving (AD), this synergy has enabled Vision Language Action (VLA) models, which translate high-level multimodal understanding into driving…

Recent advancements in open-source Visual Language Models (VLMs) such as LLaVA, Qwen-VL, and Llama have catalyzed extensive research on their integration with diverse systems. The internet-scale general knowledge encapsulated within these…

Robotics · Computer Science 2025-07-03 Cristian Gariboldi , Hayato Tokida , Ken Kinjo , Yuki Asada , Alexander Carballo

The integration of Vision-Language-Action (VLA) models into autonomous driving systems offers a unified framework for interpreting complex scenes and executing control commands. However, the necessity to incorporate historical multi-view…

Robotics · Computer Science 2026-03-30 Yiru Wang , Anqing Jiang , Shuo Wang , Yuwen Heng , Zichong Gu , Hao Sun

End-to-end autonomous driving methods aim to directly map raw sensor inputs to future driving actions such as planned trajectories, bypassing traditional modular pipelines. While these approaches have shown promise, they often operate under…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Bozhou Zhang , Nan Song , Jingyu Li , Xiatian Zhu , Jiankang Deng , Li Zhang

Autonomous driving visual question answering (AD-VQA) aims to answer questions related to perception, prediction, and planning based on given driving scene images, heavily relying on the model's spatial understanding capabilities. Prior…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Zhiyuan Zhang , Xiaofan Li , Zhihao Xu , Wenjie Peng , Zijian Zhou , Miaojing Shi , Shuangping Huang

Vision-Language-Action (VLA) models offer a promising autonomous driving paradigm for leveraging world knowledge and reasoning capabilities, especially in long-tail scenarios. However, existing VLA models often struggle with the high…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Zewei Zhou , Ruining Yang , Xuewei , Qi , Yiluan Guo , Sherry X. Chen , Tao Feng , Kateryna Pistunova , Yishan Shen , Lili Su , Jiaqi Ma

Autonomous driving requires rich contextual comprehension and precise predictive reasoning to navigate dynamic and complex environments safely. Vision-Language Models (VLMs) and Driving World Models (DWMs) have independently emerged as…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Jingyu Li , Bozhou Zhang , Xin Jin , Jiankang Deng , Xiatian Zhu , Li Zhang

Current Vision-Language-Action (VLA) models primarily focus on mapping 2D observations to actions, but exhibit notable limitations in spatiotemporal perception and reasoning: 1) spatial representations often rely on additional sensors,…

Robotics · Computer Science 2026-05-07 Wei Li , Jizhihui Liu , Li Yixing , Junwen Tong , Rui Shao , Liqiang Nie

Video content comprehension is essential for various applications, ranging from video analysis to interactive systems. Despite advancements in large-scale vision-language models (VLMs), these models often struggle to capture the nuanced,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Shuyi Zhang , Xiaoshuai Hao , Yingbo Tang , Lingfeng Zhang , Pengwei Wang , Zhongyuan Wang , Hongxuan Ma , Shanghang Zhang

Vision-language navigation (VLN) requires intelligent agents to navigate environments by interpreting linguistic instructions alongside visual observations, serving as a cornerstone task in Embodied AI. Current VLN research for unmanned…

End-to-end autonomous driving frameworks face persistent challenges in generalization, training efficiency, and interpretability. While recent methods leverage Vision-Language Models (VLMs) through supervised learning on large-scale…

Robotics · Computer Science 2025-12-11 Lin Li , Yuxin Cai , Jianwu Fang , Jianru Xue , Chen Lv

Vision-Language-Action (VLA) models have emerged as a promising paradigm for robot learning, but their representations are still largely inherited from static image-text pretraining, leaving physical dynamics to be learned from…

Robotics · Computer Science 2026-03-24 Teli Ma , Jia Zheng , Zifan Wang , Chunli Jiang , Andy Cui , Junwei Liang , Shuo Yang

Visual-Language-Action (VLA) models represent a paradigm shift in embodied AI, yet existing frameworks often struggle with imprecise spatial perception, suboptimal multimodal fusion, and instability in reinforcement learning. To bridge…

Robotics · Computer Science 2026-04-27 Haoxiang Jie , Yaoyuan Yan , Xiangyu Wei , Kailin Wang , Hongjie Yan , Zhiyou Heng , Daocheng Chen

Recent studies have explored leveraging the world knowledge and cognitive capabilities of Vision-Language Models (VLMs) to address the long-tail problem in end-to-end autonomous driving. However, existing methods typically formulate…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yongkang Li , Kaixin Xiong , Xiangyu Guo , Fang Li , Sixu Yan , Gangwei Xu , Lijun Zhou , Long Chen , Haiyang Sun , Bing Wang , Kun Ma , Guang Chen , Hangjun Ye , Wenyu Liu , Xinggang Wang

Large vision-language models (VLMs) for autonomous driving (AD) are evolving beyond perception and cognition tasks toward motion planning. However, we identify two critical challenges in this direction: (1) VLMs tend to learn shortcuts by…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Yue Li , Meng Tian , Dechang Zhu , Jiangtong Zhu , Zhenyu Lin , Zhiwei Xiong , Xinhai Zhao

Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) models capture interaction-aware temporal dependencies via causal…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Xiyang Wang , Xinlin Wang , Tingguang Zhou , Gong Chen , Xingtai Gui , Zhi Xu , Xiaolei Wu , Feiyang Tan , Hangning Zhou , Mu Yang

The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long tail scenarios. However, these models…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Nikos Theodoridis , Reenu Mohandas , Ganesh Sistu , Anthony Scanlan , Ciarán Eising , Tim Brophy

The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabilities from 2D to full 3D understanding is crucial for…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Shihao Wang , Zhiding Yu , Xiaohui Jiang , Shiyi Lan , Min Shi , Nadine Chang , Jan Kautz , Ying Li , Jose M. Alvarez

The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabilities from 2D to full 3D understanding is crucial for…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Shihao Wang , Zhiding Yu , Xiaohui Jiang , Shiyi Lan , Min Shi , Nadine Chang , Jan Kautz , Ying Li , Jose M. Alvarez
‹ Prev 1 4 5 6 7 8 10 Next ›