English
Related papers

Related papers: GenieDrive: Towards Physics-Aware Driving World Mo…

200 papers

Autonomous driving relies on robust models trained on high-quality, large-scale multi-view driving videos. While world models offer a cost-effective solution for generating realistic driving videos, they struggle to maintain instance-level…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Zhuoran Yang , Xi Guo , Chenjing Ding , Chiyu Wang , Wei Wu , Yanyong Zhang

In this paper, we introduce the first large-scale video prediction model in the autonomous driving discipline. To eliminate the restriction of high-cost data collection and empower the generalization ability of our model, we acquire massive…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Jiazhi Yang , Shenyuan Gao , Yihang Qiu , Li Chen , Tianyu Li , Bo Dai , Kashyap Chitta , Penghao Wu , Jia Zeng , Ping Luo , Jun Zhang , Andreas Geiger , Yu Qiao , Hongyang Li

World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Tianze Xia , Yongkang Li , Lijun Zhou , Jingfeng Yao , Kaixin Xiong , Haiyang Sun , Bing Wang , Kun Ma , Guang Chen , Hangjun Ye , Wenyu Liu , Xinggang Wang

The field of autonomous driving increasingly demands high-quality annotated training data. In this paper, we propose Panacea, an innovative approach to generate panoramic and controllable videos in driving scenarios, capable of yielding an…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Yuqing Wen , Yucheng Zhao , Yingfei Liu , Fan Jia , Yanhui Wang , Chong Luo , Chi Zhang , Tiancai Wang , Xiaoyan Sun , Xiangyu Zhang

Vision-centric autonomous driving has demonstrated excellent performance with economical sensors. As the fundamental step, 3D perception aims to infer 3D information from 2D images based on 3D-2D projection. This makes driving perception…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Ye Li , Wenzhao Zheng , Xiaonan Huang , Kurt Keutzer

Motion prediction of vehicles is critical but challenging due to the uncertainties in complex environments and the limited visibility caused by occlusions and limited sensor ranges. In this paper, we study a new task, safety-aware motion…

Computer Vision and Pattern Recognition · Computer Science 2021-09-06 Xuanchi Ren , Tao Yang , Li Erran Li , Alexandre Alahi , Qifeng Chen

Understanding the 3D geometry and semantics of driving scenes is critical for safe autonomous driving. Recent advances in 3D occupancy prediction have improved scene representation but often suffer from visual inconsistencies, leading to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Loïck Chambon , Eloi Zablocki , Alexandre Boulch , Mickaël Chen , Matthieu Cord

Generative AI (GenAI) is rapidly advancing the field of Autonomous Driving (AD), extending beyond traditional applications in text, image, and video generation. We explore how generative models can enhance automotive tasks, such as static…

3D occupancy, an advanced perception technology for driving scenarios, represents the entire scene without distinguishing between foreground and background by quantifying the physical space into a grid map. The widely adopted…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Jinke Li , Xiao He , Chonghua Zhou , Xiaoqiang Cheng , Yang Wen , Dan Zhang

This paper introduces a novel architecture for trajectory-conditioned forecasting of future 3D scene occupancy. In contrast to methods that rely on variational autoencoders (VAEs) to generate discrete occupancy tokens, which inherently…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Jiayuan Du , Yiming Zhao , Zhenglong Guo , Yong Pan , Wenbo Hou , Zhihui Hao , Kun Zhan , Qijun Chen

World models that forecast environmental changes from actions are vital for autonomous driving models with strong generalization. The prevailing driving world model mainly build on video prediction model. Although these models can produce…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Jingcheng Ni , Yuxin Guo , Yichen Liu , Rui Chen , Lewei Lu , Zehuan Wu

World models and video generation are pivotal technologies in the domain of autonomous driving, each playing a critical role in enhancing the robustness and reliability of autonomous systems. World models, which simulate the dynamics of…

Artificial Intelligence · Computer Science 2024-11-06 Ao Fu , Yi Zhou , Tao Zhou , Yi Yang , Bojun Gao , Qun Li , Guobin Wu , Ling Shao

We present DriveGen3D, a novel framework for generating high-quality and highly controllable dynamic 3D driving scenes that addresses critical limitations in existing methodologies. Current approaches to driving scene synthesis either…

We introduce Genie, the first generative interactive environment trained in an unsupervised manner from unlabelled Internet videos. The model can be prompted to generate an endless variety of action-controllable virtual worlds described…

Autonomous parking fundamentally differs from on-road driving due to its frequent direction changes and complex maneuvering requirements. However, existing End-to-End (E2E) planning methods often simplify the parking task into a geometric…

Robotics · Computer Science 2026-02-26 Jishu Miao , Han Chen , Jiankun Zhai , Qi Liu , Tsubasa Hirakawa , Takayoshi Yamashita , Hironobu Fujiyoshi

Reliable anticipation of traffic accidents is essential for advancing autonomous driving systems. However, this objective is limited by two fundamental challenges: the scarcity of diverse, high-quality training data and the frequent absence…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Yanchen Guan , Haicheng Liao , Chengyue Wang , Xingcheng Liu , Jiaxun Zhang , Zhenning Li

With the increasing popularity of autonomous driving based on the powerful and unified bird's-eye-view (BEV) representation, a demand for high-quality and large-scale multi-view video data with accurate annotation is urgently required.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Xiaofan Li , Yifu Zhang , Xiaoqing Ye

Images and videos are discrete 2D projections of the 4D world (3D space + time). Most visual understanding, prediction, and generation operate directly on 2D observations, leading to suboptimal performance. We propose SeeU, a novel approach…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yu Yuan , Tharindu Wickremasinghe , Zeeshan Nadir , Xijun Wang , Yiheng Chi , Stanley H. Chan

Urban scene synthesis with video generation models has recently shown great potential for autonomous driving. Existing video generation approaches to autonomous driving primarily focus on RGB video generation and lack the ability to support…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Guile Wu , David Huang , Dongfeng Bai , Bingbing Liu

Human drivers adeptly navigate complex scenarios by utilizing rich attentional semantics, but the current autonomous systems struggle to replicate this ability, as they often lose critical semantic information when converting 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Pei Liu , Haipeng Liu , Haichao Liu , Xin Liu , Jinxin Ni , Jun Ma