English
Related papers

Related papers: DriveVA: Video Action Models are Zero-Shot Drivers

200 papers

Data scaling plays a pivotal role in the pursuit of general intelligence. However, the prevailing perception-planning paradigm in autonomous driving relies heavily on expensive manual annotations to supervise trajectory planning, which…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Jiawei Xu , Zhizhou Zhong , Zhijian Shu , Mingkai Jia , Mingxiao Li , Jia-Wang Bian , Qian Zhang , Kaicheng Zhang , Jin Xie , Jian Yang , Wei Yin

Previous Vision-Language-Action models face critical limitations in navigation: scarce, diverse data from labor-intensive collection and static representations that fail to capture temporal dynamics and physical laws. We propose NavDreamer,…

Robotics · Computer Science 2026-02-11 Xijie Huang , Weiqi Gai , Tianyue Wu , Congyu Wang , Zhiyang Liu , Xin Zhou , Yuze Wu , Fei Gao

Autonomous driving requires reasoning about how the environment evolves and planning actions accordingly. Existing world-model-based approaches typically predict future scenes first and plan afterwards, resulting in open-loop imagination…

Robotics · Computer Science 2026-03-31 Qiqi Liu , Huan Xu , Jingyu Li , Bin Sun , Zhihui Hao , Dangen She , Xiatian Zhu , Li Zhang

A primary hurdle of autonomous driving in urban environments is understanding complex and long-tail scenarios, such as challenging road conditions and delicate human behaviors. We introduce DriveVLM, an autonomous driving system leveraging…

Computer Vision and Pattern Recognition · Computer Science 2024-06-26 Xiaoyu Tian , Junru Gu , Bailin Li , Yicheng Liu , Yang Wang , Zhiyong Zhao , Kun Zhan , Peng Jia , Xianpeng Lang , Hang Zhao

This work highlights that video world modeling, alongside vision-language pre-training, establishes a fresh and independent foundation for robot learning. Intuitively, video world models provide the ability to imagine the near future by…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Lin Li , Qihang Zhang , Yiming Luo , Shuai Yang , Ruilin Wang , Fei Han , Mingrui Yu , Zelin Gao , Nan Xue , Xing Zhu , Yujun Shen , Yinghao Xu

Collecting multi-view driving scenario videos to enhance the performance of 3D visual perception tasks presents significant challenges and incurs substantial costs, making generative models for realistic data an appealing alternative. Yet,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Junpeng Jiang , Gangyi Hong , Miao Zhang , Hengtong Hu , Kun Zhan , Rui Shao , Liqiang Nie

Real-world driving requires people to observe the current environment, anticipate the future, and make appropriate driving decisions. This requirement is aligned well with the capabilities of world models, which understand the environment…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Xiaodong Wang , Peixi Peng

Recent successful video generation systems that predict and create realistic automotive driving scenes from short video inputs assign tokenization, future state prediction (world model), and video decoding to dedicated models. These…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Björn Möller , Zhengyang Li , Malte Stelzer , Thomas Graave , Fabian Bettels , Muaaz Ataya , Tim Fingscheidt

Action anticipation is critical in scenarios where one needs to react before the action is finalized. This is, for instance, the case in automated driving, where a car needs to, e.g., avoid hitting pedestrians and respect traffic lights.…

Computer Vision and Pattern Recognition · Computer Science 2018-10-31 Mohammad Sadegh Aliakbarian , Fatemeh Sadat Saleh , Mathieu Salzmann , Basura Fernando , Lars Petersson , Lars Andersson

We propose DynVLA, a driving VLA model that introduces a new CoT paradigm termed Dynamics CoT. DynVLA forecasts compact world dynamics before action generation, enabling more informed and physically grounded decision-making. To obtain…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Shuyao Shang , Bing Zhan , Yunfei Yan , Yuqi Wang , Yingyan Li , Yasong An , Xiaoman Wang , Jierui Liu , Lu Hou , Lue Fan , Zhaoxiang Zhang , Tieniu Tan

End-to-end autonomous driving models based on Vision-Language-Action (VLA) architectures have shown promising results by learning driving policies through behavior cloning on expert demonstrations. However, imitation learning inherently…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Zihao Sheng , Xin Ye , Jingru Luo , Sikai Chen , Liu Ren

Vision-centric autonomous driving has recently raised wide attention due to its lower cost. Pre-training is essential for extracting a universal representation. However, current vision-centric pre-training typically relies on either 2D or…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Chen Min , Dawei Zhao , Liang Xiao , Jian Zhao , Xinli Xu , Zheng Zhu , Lei Jin , Jianshu Li , Yulan Guo , Junliang Xing , Liping Jing , Yiming Nie , Bin Dai

World models, especially in autonomous driving, are trending and drawing extensive attention due to their capacity for comprehending driving environments. The established world model holds immense potential for the generation of…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Xiaofeng Wang , Zheng Zhu , Guan Huang , Xinze Chen , Jiagang Zhu , Jiwen Lu

Learning a human-like driving policy from large-scale driving demonstrations is promising, but the uncertainty and non-deterministic nature of planning make it challenging. Existing learning-based planning methods follow a deterministic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Bo Jiang , Shaoyu Chen , Hao Gao , Bencheng Liao , Qian Zhang , Wenyu Liu , Xinggang Wang

Camera sensor simulation serves as a critical role for autonomous driving (AD), e.g. evaluating vision-based AD algorithms. While existing approaches have leveraged generative models for controllable image/video generation, they remain…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Wenchao Sun , Xuewu Lin , Keyu Chen , Zixiang Pei , Yining Shi , Chuang Zhang , Sifa Zheng

Autonomous driving heavily relies on accurate and robust spatial perception. Many failures arise from inaccuracies and instability, especially in long-tail scenarios and complex interactions. However, current vision-language models are weak…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Jianhua Han , Meng Tian , Jiangtong Zhu , Fan He , Huixin Zhang , Sitong Guo , Dechang Zhu , Hao Tang , Pei Xu , Yuze Guo , Minzhe Niu , Haojie Zhu , Qichao Dong , Xuechao Yan , Siyuan Dong , Lu Hou , Qingqiu Huang , Xiaosong Jia , Hang Xu

Comprehensive situational awareness is essential for autonomous vehicles operating in safety-critical environments, as it enables the identification and mitigation of potential risks. Although recent Multimodal Large Language Models (MLLMs)…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Sainithin Artham , Shankar Gangisetty , Avijit Dasgupta , C. V. Jawahar

Vision-Language-Action (VLA) models have advanced autonomous driving, but existing benchmarks still lack scenario diversity, reliable action-level annotation, and evaluation protocols aligned with human preferences. To address these…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Yuhan Hao , Zhengning Li , Lei Sun , Weilong Wang , Naixin Yi , Sheng Song , Caihong Qin , Mofan Zhou , Yifei Zhan , Xianpeng Lang

Existing latent world models for autonomous driving have opened a promising path toward future-aware driving intelligence. However, they typically treat future latent states as prediction targets or auxiliary signals, rather than directly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Yufeng Hong , Xiaotian Zhou , Yingyan Li , Xiangpo Zhou , Lin Liu , Yadan Luo , Shaoqing Xu , Lei Yang , Ziying Song

Autonomous driving has long relied on modular "Perception-Decision-Action" pipelines, where hand-crafted interfaces and rule-based components often break down in complex or long-tailed scenarios. Their cascaded design further propagates…