English
Related papers

Related papers: DriveVA: Video Action Models are Zero-Shot Drivers

200 papers

Generating high-fidelity, temporally consistent videos in autonomous driving scenarios faces a significant challenge, e.g. problematic maneuvers in corner cases. Despite recent video generation works are proposed to tackcle the mentioned…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Junpeng Jiang , Gangyi Hong , Lijun Zhou , Enhui Ma , Hengtong Hu , Xia Zhou , Jie Xiang , Fan Liu , Kaicheng Yu , Haiyang Sun , Kun Zhan , Peng Jia , Miao Zhang

Driver visual attention prediction is a critical task in autonomous driving and human-computer interaction (HCI) research. Most prior studies focus on estimating attention allocation at a single moment in time, typically using static RGB…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Kaiser Hamid , Khandakar Ashrafi Akbar , Nade Liang

Efficient trajectory planning for urban intersections is currently one of the most challenging tasks for an Autonomous Vehicle (AV). Courteous behavior towards other traffic participants, the AV's comfort and its progression in the…

Robotics · Computer Science 2020-10-08 Oliver Speidel , Maximilian Graf , Ankit Kaushik , Thanh Phan-Huu , Andreas Wedel , Klaus Dietmayer

Modern vehicles equip dashcams that primarily collect visual evidence for traffic accidents. However, most of the video data collected by dashcams that is not related to traffic accidents is discarded without any use. In this paper, we…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-12-02 Seyul Lee , Jayden King , Young Choon Lee , Hyuck Han , Sooyong Kang

In this paper, we introduce the first large-scale video prediction model in the autonomous driving discipline. To eliminate the restriction of high-cost data collection and empower the generalization ability of our model, we acquire massive…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Jiazhi Yang , Shenyuan Gao , Yihang Qiu , Li Chen , Tianyu Li , Bo Dai , Kashyap Chitta , Penghao Wu , Jia Zeng , Ping Luo , Jun Zhang , Andreas Geiger , Yu Qiao , Hongyang Li

Physics-aware driving world model is essential for drive planning, out-of-distribution data synthesis, and closed-loop evaluation. However, existing methods often rely on a single diffusion model to directly map driving actions to videos,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Zhenya Yang , Zhe Liu , Yuxiang Lu , Liping Hou , Chenxuan Miao , Siyi Peng , Bailan Feng , Xiang Bai , Hengshuang Zhao

Navigating unfamiliar environments presents significant challenges for household robots, requiring the ability to recognize and reason about novel decoration and layout. Existing reinforcement learning methods cannot be directly transferred…

Robotics · Computer Science 2025-02-20 Yiran Qin , Ao Sun , Yuze Hong , Benyou Wang , Ruimao Zhang

Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, maintain object continuity across viewpoints and time, and reason about spatial relations,…

Anticipating the motion of neighboring vehicles is crucial for autonomous driving, especially on congested highways where even slight motion variations can result in catastrophic collisions. An accurate prediction of a future trajectory…

Computer Vision and Pattern Recognition · Computer Science 2023-04-20 Fuad Hasan , Hailong Huang

If a Large Language Model (LLM) were to take a driving knowledge test today, would it pass? Beyond standard spatial and visual question-answering (QA) tasks on current autonomous driving benchmarks, driving knowledge tests require a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-01 Maolin Wei , Wanzhou Liu , Eshed Ohn-Bar

The inherent sequential modeling capabilities of autoregressive models make them a formidable baseline for end-to-end planning in autonomous driving. Nevertheless, their performance is constrained by a spatio-temporal misalignment, as the…

Robotics · Computer Science 2025-09-26 Jianbo Zhao , Taiyu Ban , Xiangjie Li , Xingtai Gui , Hangning Zhou , Lei Liu , Hongwei Zhao , Bin Li

Autonomous vehicles interacting with other traffic participants heavily rely on the perception and prediction of other agents' behaviors to plan safe trajectories. However, as occlusions limit the vehicle's perception ability, reasoning…

Robotics · Computer Science 2021-08-04 Zixu Zhang , Jaime F. Fisac

Motion forecasting for agents in autonomous driving is highly challenging due to the numerous possibilities for each agent's next action and their complex interactions in space and time. In real applications, motion forecasting takes place…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Nan Song , Bozhou Zhang , Xiatian Zhu , Li Zhang

The ability to predict future outcomes given control actions is fundamental for physical reasoning. However, such predictive models, often called world models, remains challenging to learn and are typically developed for task-specific…

Robotics · Computer Science 2025-02-04 Gaoyue Zhou , Hengkai Pan , Yann LeCun , Lerrel Pinto

Autonomous driving requires rich contextual comprehension and precise predictive reasoning to navigate dynamic and complex environments safely. Vision-Language Models (VLMs) and Driving World Models (DWMs) have independently emerged as…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Jingyu Li , Bozhou Zhang , Xin Jin , Jiankang Deng , Xiatian Zhu , Li Zhang

Autonomous Driving (AD) systems have made notable progress, but their performance in long-tail, safety-critical scenarios remains limited. These rare cases contribute a disproportionate number of accidents. Vision-Language Action (VLA)…

Robotics · Computer Science 2025-09-22 Shiyu Fang , Yiming Cui , Haoyang Liang , Chen Lv , Peng Hang , Jian Sun

Vision-Language-Action (VLA) models aim to control robots for manipulation from visual observations and natural-language instructions. However, existing hierarchical and autoregressive paradigms often introduce architectural overhead,…

Recent advancements in world models have revolutionized dynamic environment simulation, allowing systems to foresee future states and assess potential actions. In autonomous driving, these capabilities help vehicles anticipate the behavior…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Anthony Chen , Wenzhao Zheng , Yida Wang , Xueyang Zhang , Kun Zhan , Peng Jia , Kurt Keutzer , Shanghang Zhang

Recently, the diffusion model has emerged as a powerful generative technique for robotic policy learning, capable of modeling multi-mode action distributions. Leveraging its capability for end-to-end autonomous driving is a promising…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Bencheng Liao , Shaoyu Chen , Haoran Yin , Bo Jiang , Cheng Wang , Sixu Yan , Xinbang Zhang , Xiangyu Li , Ying Zhang , Qian Zhang , Xinggang Wang

The end-to-end learning ability of self-driving vehicles has achieved significant milestones over the last decade owing to rapid advances in deep learning and computer vision algorithms. However, as autonomous driving technology is a…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Shahin Atakishiyev , Mohammad Salameh , Housam Babiker , Randy Goebel