English
Related papers

Related papers: DLWM: Dual Latent World Models enable Holistic Gau…

200 papers

Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, without the ability to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Tianchen Deng , Xuefeng Chen , Yi Chen , Qu Chen , Yuyao Xu , Lijin Yang , Le Xu , Yu Zhang , Bo Zhang , Wuxiong Huang , Hesheng Wang

Vision-based autonomous driving shows great potential due to its satisfactory performance and low costs. Most existing methods adopt dense representations (e.g., bird's eye view) or sparse representations (e.g., instance boxes) for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Wenzhao Zheng , Junjie Wu , Yao Zheng , Sicheng Zuo , Zixun Xie , Longchao Yang , Yong Pan , Zhihui Hao , Peng Jia , Xianpeng Lang , Shanghang Zhang

We introduce Latent-WAM, an efficient end-to-end autonomous driving framework that achieves strong trajectory planning through spatially-aware and dynamics-informed latent world representations. Existing world-model-based planners suffer…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Linbo Wang , Yupeng Zheng , Qiang Chen , Shiwei Li , Yichen Zhang , Zebin Xing , Qichao Zhang , Xiang Li , Deheng Qian , Pengxuan Yang , Yihang Dong , Ce Hao , Xiaoqing Ye , Junyu han , Yifeng Pan , Dongbin Zhao

Future 3D semantic occupancy forecasting and motion planning are central to autonomous driving, as they require models to reason about how surrounding scenes evolve and how the ego vehicle should act. Existing occupancy world models…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Cheng Chen , Hao Huang , Saurabh Bagchi

The Driving World Model (DWM), which focuses on predicting scene evolution during the driving process, has emerged as a promising paradigm in the pursuit of autonomous driving (AD). DWMs enable AD systems to better perceive, understand, and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Sifan Tu , Xin Zhou , Dingkang Liang , Xingyu Jiang , Yumeng Zhang , Xiaofan Li , Xiang Bai

Training robot policies within a learned world model is trending due to the inefficiency of real-world interactions. The established image-based world models and policies have shown prior success, but lack robust geometric information that…

Robotics · Computer Science 2025-09-18 Guanxing Lu , Baoxiong Jia , Puhao Li , Yixin Chen , Ziwei Wang , Yansong Tang , Siyuan Huang

Vision-centric autonomous driving has recently raised wide attention due to its lower cost. Pre-training is essential for extracting a universal representation. However, current vision-centric pre-training typically relies on either 2D or…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Chen Min , Dawei Zhao , Liang Xiao , Jian Zhao , Xinli Xu , Zheng Zhu , Lei Jin , Jianshu Li , Yulan Guo , Junliang Xing , Liping Jing , Yiming Nie , Bin Dai

End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual reasoning to enhance the robustness and accuracy of driving decisions. However, the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Lingjun Zhang , Changjie Wu , Linzhe Shi , Jiangyang Li , Jiaxin Liu , Lei Yang , Hang Zhang , Mu Xu , Hong Wang

3D occupancy prediction is important for autonomous driving due to its comprehensive perception of the surroundings. To incorporate sequential inputs, most existing methods fuse representations from previous frames to infer the current 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Sicheng Zuo , Wenzhao Zheng , Yuanhui Huang , Jie Zhou , Jiwen Lu

We present LatentAM, an online 3D Gaussian Splatting (3DGS) mapping framework that builds scalable latent feature maps from streaming RGB-D observations for open-vocabulary robotic perception. Instead of distilling high-dimensional…

Robotics · Computer Science 2026-02-16 Junwoon Lee , Yulun Tian

Understanding world dynamics is crucial for planning in autonomous driving. Recent methods attempt to achieve this by learning a 3D occupancy world model that forecasts future surrounding scenes based on current observation. However, 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-02-12 Xiang Li , Pengfei Li , Yupeng Zheng , Wei Sun , Yan Wang , Yilun Chen

Safe L2/L3 driving automation requires anticipating human-in-the-loop reactions during shared-control transitions. While most driving world models forecast the external environment, in-cabin intelligence remains strictly…

Robotics · Computer Science 2026-05-07 Haozhuang Chi , Daosheng Qiu , Hao Su , Haochen Liu , Zirui Li , Haoruo Zhang , Chen Lv

End-to-end autonomous driving systems increasingly rely on vision-centric world models to understand and predict their environment. However, a common ineffectiveness in these models is the full reconstruction of future scenes, which expends…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Jianbiao Mei , Yu Yang , Xuemeng Yang , Licheng Wen , Jiajun Lv , Botian Shi , Yong Liu

This paper introduces a novel method for open-vocabulary 3D scene querying in autonomous driving by combining Language Embedded 3D Gaussians with Large Language Models (LLMs). We propose utilizing LLMs to generate both contextually…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Amirhosein Chahe , Lifeng Zhou

Autonomous driving heavily relies on accurate and robust spatial perception. Many failures arise from inaccuracies and instability, especially in long-tail scenarios and complex interactions. However, current vision-language models are weak…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Jianhua Han , Meng Tian , Jiangtong Zhu , Fan He , Huixin Zhang , Sitong Guo , Dechang Zhu , Hao Tang , Pei Xu , Yuze Guo , Minzhe Niu , Haojie Zhu , Qichao Dong , Xuechao Yan , Siyuan Dong , Lu Hou , Qingqiu Huang , Xiaosong Jia , Hang Xu

The rise of multi-modal large language models(MLLMs) has spurred their applications in autonomous driving. Recent MLLM-based methods perform action by learning a direct mapping from perception to action, neglecting the dynamics of the world…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Julong Wei , Shanshuai Yuan , Pengfei Li , Qingda Hu , Zhongxue Gan , Wenchao Ding

World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Tianze Xia , Yongkang Li , Lijun Zhou , Jingfeng Yao , Kaixin Xiong , Haiyang Sun , Bing Wang , Kun Ma , Guang Chen , Hangjun Ye , Wenyu Liu , Xinggang Wang

In autonomous driving, predicting future events in advance and evaluating the foreseeable risks empowers autonomous vehicles to better plan their actions, enhancing safety and efficiency on the road. To this end, we propose Drive-WM, the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Yuqi Wang , Jiawei He , Lue Fan , Hongxin Li , Yuntao Chen , Zhaoxiang Zhang

Vision-based deep learning (DL) methods have made great progress in learning autonomous driving models from large-scale crowd-sourced video datasets. They are trained to predict instantaneous driving behaviors from video data captured by…

Human-Computer Interaction · Computer Science 2021-09-24 Suphanut Jamonnak , Ye Zhao , Xinyi Huang , Md Amiruzzaman

Recent breakthroughs in autonomous driving have been propelled by advances in robust world modeling, fundamentally transforming how vehicles interpret dynamic scenes and execute safe decision-making. World models have emerged as a linchpin…

Robotics · Computer Science 2025-09-11 Tuo Feng , Wenguan Wang , Yi Yang
‹ Prev 1 2 3 10 Next ›