中文
相关论文

相关论文: DriveVA: Video Action Models are Zero-Shot Drivers

200 篇论文

World models and video generation are pivotal technologies in the domain of autonomous driving, each playing a critical role in enhancing the robustness and reliability of autonomous systems. World models, which simulate the dynamics of…

人工智能 · 计算机科学 2024-11-06 Ao Fu , Yi Zhou , Tao Zhou , Yi Yang , Bojun Gao , Qun Li , Guobin Wu , Ling Shao

Dynamic scene reconstruction in autonomous driving remains a fundamental challenge due to significant temporal variations, moving objects, and complex scene dynamics. Existing feed-forward 3D models have demonstrated strong performance in…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Zhuolin He , Jing Li , Guanghao Li , Xiaolei Chen , Jiacheng Tang , Siyang Zhang , Zhounan Jin , Feipeng Cai , Bin Li , Jian Pu , Jia Cai , Xiangyang Xue

World foundation models aim to simulate the evolution of the real world with physically plausible behavior. Unlike prior methods that handle spatial and temporal correlations separately, we propose RAYNOVA, a geometry-agonistic multiview…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Yichen Xie , Chensheng Peng , Mazen Abdelfattah , Yihan Hu , Jiezhi Yang , Eric Higgins , Ryan Brigden , Masayoshi Tomizuka , Wei Zhan

Autonomous vehicles (AVs) are poised to revolutionize global transportation systems. However, its widespread acceptance and market penetration remain significantly below expectations. This gap is primarily driven by persistent challenges in…

机器人学 · 计算机科学 2026-02-10 Yuxin Zhang , Cheng Wang , Hubert P. H. Shum

We present WorldVLA, an autoregressive action world model that unifies action and image understanding and generation. Our WorldVLA intergrates Vision-Language-Action (VLA) model and world model in one single framework. The world model…

机器人学 · 计算机科学 2025-06-27 Jun Cen , Chaohui Yu , Hangjie Yuan , Yuming Jiang , Siteng Huang , Jiayan Guo , Xin Li , Yibing Song , Hao Luo , Fan Wang , Deli Zhao , Hao Chen

Search-based testing is critical for evaluating the safety and reliability of autonomous driving systems (ADSs). However, existing approaches are often built on heterogeneous frameworks (e.g., distinct scenario spaces, simulators, and…

软件工程 · 计算机科学 2026-01-12 Mingfei Cheng , Lionel Briand , Yuan Zhou

When uncertainty is high, self-driving vehicles may halt for safety and benefit from the access to remote human operators who can provide high-level guidance. This paradigm, known as {shared autonomy}, enables autonomous vehicle and remote…

机器人学 · 计算机科学 2025-11-18 Shuangyu Xie , Kaiyuan Chen , Wenjing Chen , Chengyuan Qian , Christian Juette , Liu Ren , Dezhen Song , Ken Goldberg

Realtime 4D reconstruction for dynamic scenes remains a crucial challenge for autonomous driving perception. Most existing methods rely on depth estimation through self-supervision or multi-modality sensor fusion. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Xin Fei , Wenzhao Zheng , Yueqi Duan , Wei Zhan , Masayoshi Tomizuka , Kurt Keutzer , Jiwen Lu

Self-driving vehicles (SDVs) hold great potential for improving traffic safety and are poised to positively affect the quality of life of millions of people. To unlock this potential one of the critical aspects of the autonomous technology…

机器人学 · 计算机科学 2020-10-27 Henggang Cui , Thi Nguyen , Fang-Chieh Chou , Tsung-Han Lin , Jeff Schneider , David Bradley , Nemanja Djuric

Reconstructing large-scale dynamic driving scenes remains challenging due to the coexistence of static environments with extreme depth variation and diverse dynamic actors exhibiting complex motions. Existing Gaussian Splatting based…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Cong Wang , Ruiqi Song , Wei Tian , Chenming Zhang , Lingxi Li , Long Chen

Fusing sensors with complementary modalities is crucial for maintaining a stable and comprehensive understanding of abnormal driving scenes. However, Multimodal Large Language Models (MLLMs) are underexplored for leveraging multi-sensor…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Mingzhe Tao , Ruiping Liu , Junwei Zheng , Yufan Chen , Kedi Ying , M. Saquib Sarfraz , Kailun Yang , Jiaming Zhang , Rainer Stiefelhagen

Autonomous driving world models are expected to work effectively across three core dimensions: state, action, and reward. Existing models, however, are typically restricted to limited state modalities, short video sequences, imprecise…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Bohan Li , Zhuang Ma , Dalong Du , Baorui Peng , Zhujin Liang , Zhenqiang Liu , Chao Ma , Yueming Jin , Hao Zhao , Wenjun Zeng , Xin Jin

Vision-Language-Action (VLA) models represent a pivotal advance in embodied intelligence, yet they confront critical barriers to real-world deployment, most notably catastrophic forgetting. This issue stems from their overreliance on…

机器人学 · 计算机科学 2026-05-12 Zheng Huang , Mingyu Liu , Xiaoyi Lin , Muzhi Zhu , Canyu Zhao , Zongze Du , Ye Lin , Xiaoman Li , Yiduo Jia , Hao Zhong , Hao Chen , Chunhua Shen

A generalist robot should perform effectively across various environments. However, most existing approaches heavily rely on scaling action-annotated data to enhance their capabilities. Consequently, they are often limited to single…

机器人学 · 计算机科学 2025-11-04 Qingwen Bu , Yanting Yang , Jisong Cai , Shenyuan Gao , Guanghui Ren , Maoqing Yao , Ping Luo , Hongyang Li

With the rise of vision-language models (VLM), their application for autonomous driving (VLM4AD) has gained significant attention. Meanwhile, in autonomous driving, closed-loop evaluation has become widely recognized as a more reliable…

机器人学 · 计算机科学 2026-04-03 Xiaosong Jia , Yuqian Shao , Zhenjie Yang , Qifeng Li , Zhiyuan Zhang , Junchi Yan

Feedforward reconstruction is crucial for autonomous driving applications, where rapid scene reconstruction enables efficient utilization of large-scale driving datasets in closed-loop simulation and other downstream tasks, eliminating the…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Zhongrui Yu , Zhao Wang , Yijia Xie , Yida Wang , Xueyang Zhang , Yifei Zhan , Kun Zhan

In recent years, the field of autonomous driving has attracted increasingly significant public interest. Accurately forecasting the future behavior of various traffic participants is essential for the decision-making of Autonomous Vehicles…

机器人学 · 计算机科学 2025-02-14 Jianxin Shi , Jinhao Chen , Yuandong Wang , Li Sun , Chunyang Liu , Wei Xiong , Tianyu Wo

Collaborative navigation becomes essential in situations of occluded scenarios in autonomous driving where independent driving policies are likely to lead to collisions. One promising approach to address this issue is through the use of…

机器人学 · 计算机科学 2024-12-12 Leandro Parada , Hanlin Tian , Jose Escribano , Panagiotis Angeloudis

Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language pre-training. This progress can be misleading. Despite the zero-shot perception and…

Vision-Language-Action (VLA) models offer a promising autonomous driving paradigm for leveraging world knowledge and reasoning capabilities, especially in long-tail scenarios. However, existing VLA models often struggle with the high…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Zewei Zhou , Ruining Yang , Xuewei , Qi , Yiluan Guo , Sherry X. Chen , Tao Feng , Kateryna Pistunova , Yishan Shen , Lili Su , Jiaqi Ma