English
Related papers

Related papers: DriveVA: Video Action Models are Zero-Shot Drivers

200 papers

A unified video and action model holds significant promise for robotics, where videos provide rich scene information for action prediction, and actions provide dynamics information for video prediction. However, effectively combining video…

Robotics · Computer Science 2025-04-28 Shuang Li , Yihuai Gao , Dorsa Sadigh , Shuran Song

Trajectory prediction for traffic agents is critical for safe autonomous driving. However, achieving effective zero-shot generalization in previously unseen domains remains a significant challenge. Motivated by the consistent nature of…

Artificial Intelligence · Computer Science 2026-02-17 Zhenyu Zong , Yuchen Wang , Haohong Lin , Lu Gan , Huajie Shao

In autonomous driving, dynamic environment and corner cases pose significant challenges to the robustness of ego vehicle's state understanding and decision making. We introduce VDRive, a novel pipeline for end-to-end autonomous driving that…

Robotics · Computer Science 2026-02-11 Ziang Guo , Zufeng Zhang

Recent video diffusion models generate photorealistic, temporally coherent videos, yet they fall short as reliable world models for autonomous driving, where structured motion and physically consistent interactions are essential. Adapting…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Ahmad Rahimi , Valentin Gerard , Eloi Zablocki , Matthieu Cord , Alexandre Alahi

Driving world models are used to simulate futures by video generation based on the condition of the current state and actions. However, current models often suffer serious error accumulations when predicting the long-term future, which…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Xiaodong Wang , Zhirong Wu , Peixi Peng

Vision-based deep learning (DL) methods have made great progress in learning autonomous driving models from large-scale crowd-sourced video datasets. They are trained to predict instantaneous driving behaviors from video data captured by…

Human-Computer Interaction · Computer Science 2021-09-24 Suphanut Jamonnak , Ye Zhao , Xinyi Huang , Md Amiruzzaman

Nonprehensile manipulation is crucial for handling objects that are too thin, large, or otherwise ungraspable in unstructured environments. While conventional planning-based approaches struggle with complex contact modeling, learning-based…

Robotics · Computer Science 2025-07-28 Jiangran Lyu , Ziming Li , Xuesong Shi , Chaoyi Xu , Yizhou Wang , He Wang

End-to-End (E2E) solutions have emerged as a mainstream approach for autonomous driving systems, with Vision-Language-Action (VLA) models representing a new paradigm that leverages pre-trained multimodal knowledge from Vision-Language…

Robotics · Computer Science 2025-09-25 Pengxiang Li , Yinan Zheng , Yue Wang , Huimin Wang , Hang Zhao , Jingjing Liu , Xianyuan Zhan , Kun Zhan , Xianpeng Lang

One of the challenges in vision-based driving trajectory generation is dealing with out-of-distribution scenarios. In this paper, we propose a domain generalization method for vision-based driving trajectory generation for autonomous…

Computer Vision and Pattern Recognition · Computer Science 2021-09-29 Yunkai Wang , Dongkun Zhang , Yuxiang Cui , Zexi Chen , Wei Jing , Junbo Chen , Rong Xiong , Yue Wang

Autonomous vehicles (AVs) rely on accurate trajectory prediction of surrounding vehicles to ensure the safety of both passengers and other road users. Trajectory prediction spans both short-term and long-term horizons, each requiring…

Robotics · Computer Science 2024-12-31 Chengyue Wang , Haicheng Liao , Kaiqun Zhu , Guohui Zhang , Zhenning Li

Autonomous driving relies on robust models trained on high-quality, large-scale multi-view driving videos. While world models offer a cost-effective solution for generating realistic driving videos, they struggle to maintain instance-level…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Zhuoran Yang , Xi Guo , Chenjing Ding , Chiyu Wang , Wei Wu , Yanyong Zhang

Generating 3D vehicle assets from in-the-wild observations is crucial to autonomous driving. Existing image-to-3D methods cannot well address this problem because they learn generation merely from image RGB information without a deeper…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Yibo Liu , Zheyuan Yang , Guile Wu , Yuan Ren , Kejian Lin , Bingbing Liu , Yang Liu , Jinjun Shan

Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) models capture interaction-aware temporal dependencies via causal…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Xiyang Wang , Xinlin Wang , Tingguang Zhou , Gong Chen , Xingtai Gui , Zhi Xu , Xiaolei Wu , Feiyang Tan , Hangning Zhou , Mu Yang

End-to-End (E2E) autonomous driving models are usually trained and evaluated with a fixed ego-vehicle, even though their driving policy is implicitly tied to vehicle dynamics. When such a model is deployed on a vehicle with different size,…

Robotics · Computer Science 2026-04-15 Haesung Oh , Jaeheung Park

Generalization under distribution shift remains a central bottleneck for closed-loop autonomous driving. Although simulators like CARLA enable safe and scalable testing, existing benchmarks rarely measure true generalization: they typically…

Robotics · Computer Science 2026-04-10 Simon Gerstenecker , Andreas Geiger , Katrin Renz

Although multi-modal large language models (MLLMs) have shown strong capabilities across diverse domains, their application in generating fine-grained 3D perception and prediction outputs in autonomous driving remains underexplored. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Zhe Liu , Runhui Huang , Rui Yang , Siming Yan , Zining Wang , Lu Hou , Di Lin , Xiang Bai , Hengshuang Zhao

Despite remarkable progress in Vision-Language-Action models (VLAs) for robot manipulation, these large pre-trained models require fine-tuning to be deployed in specific environments. These fine-tuned models are highly sensitive to camera…

Robotics · Computer Science 2026-03-09 Hyeongjun Heo , Seungyeon Woo , Sang Min Kim , Junho Kim , Junho Lee , Yonghyeon Lee , Young Min Kim

For end-to-end autonomous driving (E2E-AD), the evaluation system remains an open problem. Existing closed-loop evaluation protocols usually rely on simulators like CARLA being less realistic; while NAVSIM using real-world vision data, yet…

Robotics · Computer Science 2024-12-16 Junqi You , Xiaosong Jia , Zhiyuan Zhang , Yutao Zhu , Junchi Yan

Benchmarking vision-based driving policies is challenging. On one hand, open-loop evaluation with real data is easy, but these results do not reflect closed-loop performance. On the other, closed-loop evaluation is possible in simulation,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Daniel Dauner , Marcel Hallgarten , Tianyu Li , Xinshuo Weng , Zhiyu Huang , Zetong Yang , Hongyang Li , Igor Gilitschenski , Boris Ivanovic , Marco Pavone , Andreas Geiger , Kashyap Chitta

Visual Question Answering (VQA) models, which fall under the category of vision-language models, conventionally execute multiple downsampling processes on image inputs to strike a balance between computational efficiency and model…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Xirui Zhou , Lianlei Shan , Xiaolin Gui
‹ Prev 1 3 4 5 6 7 10 Next ›