English
Related papers

Related papers: DriveDreamer4D: World Models Are Effective Data Ma…

200 papers

Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via architectural modifications, they often incur high…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Weijie Wang , Xiaoxuan He , Youping Gu , Yifan Yang , Zeyu Zhang , Yefei He , Yanbo Ding , Xirui Hu , Donny Y. Chen , Zhiyuan He , Yuqing Yang , Bohan Zhuang

3D reconstruction and novel view synthesis are critical for validating autonomous driving systems and training advanced perception models. Recent self-supervised methods have gained significant attention due to their cost-effectiveness and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Xiao Tang , Guirong Zhuo , Cong Wang , Boyuan Zheng , Minqing Huang , Lianqing Zheng , Long Chen , Shouyi Lu

Comprehensive testing of autonomous systems through simulation is essential to ensure the safety of autonomous driving vehicles. This requires the generation of safety-critical scenarios that extend beyond the limitations of real-world data…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Hao Li , Chenming Wu , Ming Yuan , Yan Zhang , Chen Zhao , Chunyu Song , Haocheng Feng , Errui Ding , Dingwen Zhang , Jingdong Wang

We tackle the problem of producing realistic simulations of LiDAR point clouds, the sensor of preference for most self-driving vehicles. We argue that, by leveraging real data, we can simulate the complex world more realistically compared…

Computer Vision and Pattern Recognition · Computer Science 2020-06-17 Sivabalan Manivasagam , Shenlong Wang , Kelvin Wong , Wenyuan Zeng , Mikita Sazanovich , Shuhan Tan , Bin Yang , Wei-Chiu Ma , Raquel Urtasun

The landscape of video generation is shifting, from a focus on generating visually appealing clips to building virtual environments that support interaction and maintain physical plausibility. These developments point toward the emergence…

Artificial Intelligence · Computer Science 2026-02-09 Jingtong Yue , Ziqi Huang , Zhaoxi Chen , Xintao Wang , Pengfei Wan , Ziwei Liu

Realistic and diverse traffic scenarios in large quantities are crucial for the development and validation of autonomous driving systems. However, owing to numerous difficulties in the data collection process and the reliance on intensive…

Robotics · Computer Science 2025-10-07 Shuo Sun , Zekai Gu , Tianchen Sun , Jiawei Sun , Chengran Yuan , Yuhang Han , Dongen Li , Marcelo H. Ang

Large-scale scene data is essential for training and testing in robot learning. Neural reconstruction methods have promised the capability of reconstructing large physically-grounded outdoor scenes from captured sensor data. However, these…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Julian Ost , Andrea Ramazzina , Amogh Joshi , Maximilian Bömer , Mario Bijelic , Felix Heide

Current generative models struggle to synthesize dynamic 4D driving scenes that simultaneously support temporal extrapolation and spatial novel view synthesis (NVS) without per-scene optimization. A key challenge lies in finding an…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Jiazhe Guo , Yikang Ding , Xiwu Chen , Shuo Chen , Bohan Li , Yingshuang Zou , Xiaoyang Lyu , Feiyang Tan , Xiaojuan Qi , Zhiheng Li , Hao Zhao

Latent World Models enhance scene representation through temporal self-supervised learning, presenting a perception annotation-free paradigm for end-to-end autonomous driving. However, the reconstruction-oriented representation learning…

In autonomous driving, end-to-end planners directly utilize raw sensor data, enabling them to extract richer scene features and reduce information loss compared to traditional planners. This raises a crucial research question: how can we…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Yingyan Li , Lue Fan , Jiawei He , Yuqi Wang , Yuntao Chen , Zhaoxiang Zhang , Tieniu Tan

This study seeks to automate camera movement control for filming existing subjects into attractive videos, contrasting with the creation of non-existent content by directly generating the pixels. We select drone videos as our test case due…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Yunzhong Hou , Liang Zheng , Philip Torr

Self-driving vehicles (SDVs) must be rigorously tested on a wide range of scenarios to ensure safe deployment. The industry typically relies on closed-loop simulation to evaluate how the SDV interacts on a corpus of synthetic and real…

Robotics · Computer Science 2023-11-03 Jay Sarva , Jingkang Wang , James Tu , Yuwen Xiong , Sivabalan Manivasagam , Raquel Urtasun

This report presents a unified technical system addressing the two core capabilities of world models for autonomous driving: world representation and world generation. For world representation, we propose WorldRec, a feed-forward…

World Generation Models are emerging as a cornerstone of next-generation multimodal intelligence systems. Unlike traditional 2D visual generation, World Models aim to construct realistic, dynamic, and physically consistent 3D/4D worlds from…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yiting Lu , Wei Luo , Peiyan Tu , Haoran Li , Hanxin Zhu , Zihao Yu , Xingrui Wang , Xinyi Chen , Xinge Peng , Xin Li , Zhibo Chen

Generative video models, a leading approach to world modeling, face fundamental limitations. They often violate physical and logical rules, lack interactivity, and operate as opaque black boxes ill-suited for building structured, queryable…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Felix O'Mahony , Roberto Cipolla , Ayush Tewari

This paper presents an effective solution for view extrapolation in autonomous driving scenarios. Recent approaches focus on generating shifted novel view images from given viewpoints using diffusion models. However, these methods heavily…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Yuang Jia , Jinlong Wang , Jiayi Zhao , Chunlam Li , Shunzhou Wang , Wei Gao

Photorealistic 4D reconstruction of street scenes is essential for developing real-world simulators in autonomous driving. However, most existing methods perform this task offline and rely on time-consuming iterative processes, limiting…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Hao Lu , Tianshuo Xu , Wenzhao Zheng , Yunpeng Zhang , Wei Zhan , Dalong Du , Masayoshi Tomizuka , Kurt Keutzer , Yingcong Chen

4D radars, which provide 3D point cloud data along with Doppler velocity, are attractive components of modern automated driving systems due to their low cost and robustness under adverse weather conditions. However, they provide a…

Robotics · Computer Science 2026-03-13 Siqi Pei , Andras Palffy , Dariu M. Gavrila

The paradigm of learning-based robotics holds immense promise, yet its translation to real-world applications is critically hindered by the sample inefficiency and brittleness of conventional model-free reinforcement learning algorithms. In…

Robotics · Computer Science 2025-12-02 Agniprabha Chakraborty

While video-generation-based embodied world models have gained increasing attention, their reliance on large-scale embodied interaction data remains a key bottleneck. The scarcity, difficulty of collection, and high dimensionality of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Hao Li , Qiao Sun