中文
相关论文

相关论文: RenderWorld: World Model with Self-Supervised 3D L…

200 篇论文

Modeling and rendering dynamic urban driving scenes is crucial for self-driving simulation. Current high-quality methods typically rely on costly manual object tracklet annotations, while self-supervised approaches fail to capture dynamic…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Jiawei Xu , Kai Deng , Zexin Fan , Shenlong Wang , Jin Xie , Jian Yang

Referring 3D Gaussian Splatting (R3DGS), which utilizes natural language for 3D object segmentation, has emerged as a crucial capability for embodied AI. However, existing methods typically rely on expensive per-scene manual annotation and…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Yuyang Tan , Renhe Zhang , Hang Zhang , Ao Li , Xin Tan

World models have gained significant attention as a promising approach for autonomous driving. By emulating human-like perception and decision-making processes, these models can predict and adapt to dynamic environments. Existing methods…

机器人学 · 计算机科学 2025-12-03 Huiqian Li , Wei Pan , Haodong Zhang , Jin Huang , Zhihua Zhong

In recent years, end-to-end autonomous driving architectures have gained increasing attention due to their advantage in avoiding error accumulation. Most existing end-to-end autonomous driving methods are based on Imitation Learning (IL),…

人工智能 · 计算机科学 2025-04-22 Yueyuan Li , Mingyang Jiang , Songan Zhang , Wei Yuan , Chunxiang Wang , Ming Yang

We propose GGS, a Generalizable Gaussian Splatting method for Autonomous Driving which can achieve realistic rendering under large viewpoint changes. Previous generalizable 3D gaussian splatting methods are limited to rendering novel views…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Huasong Han , Kaixuan Zhou , Xiaoxiao Long , Yusen Wang , Chunxia Xiao

Autonomous driving, as an agent operating in the physical world, requires the fundamental capability to build \textit{world models} that capture how the environment evolves spatiotemporally in order to support long-term planning. At the…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Haoran Zhu , Anna Choromanska

Autonomous driving requires reasoning about how the environment evolves and planning actions accordingly. Existing world-model-based approaches typically predict future scenes first and plan afterwards, resulting in open-loop imagination…

机器人学 · 计算机科学 2026-03-31 Qiqi Liu , Huan Xu , Jingyu Li , Bin Sun , Zhihui Hao , Dangen She , Xiatian Zhu , Li Zhang

Camera-based end-to-end driving neural networks bring the promise of a low-cost system that maps camera images to driving control commands. These networks are appealing because they replace laborious hand engineered building blocks but…

计算机视觉与模式识别 · 计算机科学 2020-08-11 Abdelhak Loukkal , Yves Grandvalet , Tom Drummond , You Li

We introduce Latent-WAM, an efficient end-to-end autonomous driving framework that achieves strong trajectory planning through spatially-aware and dynamics-informed latent world representations. Existing world-model-based planners suffer…

Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, without the ability to…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Tianchen Deng , Xuefeng Chen , Yi Chen , Qu Chen , Yuyao Xu , Lijin Yang , Le Xu , Yu Zhang , Bo Zhang , Wuxiong Huang , Hesheng Wang

3D Gaussian Splatting (3D GS) has gained popularity due to its faster rendering speed and high-quality novel view synthesis. Some researchers have explored using 3D GS for reconstructing driving scenes. However, these methods often rely on…

计算机视觉与模式识别 · 计算机科学 2024-12-23 Yixiong Huo , Guangfeng Jiang , Hongyang Wei , Ji Liu , Song Zhang , Han Liu , Xingliang Huang , Mingjie Lu , Jinzhang Peng , Dong Li , Lu Tian , Emad Barsoum

In autonomous vehicles, understanding the surrounding 3D environment of the ego vehicle in real-time is essential. A compact way to represent scenes while encoding geometric distances and semantic object information is via 3D semantic…

机器人学 · 计算机科学 2024-05-21 Samuel Sze , Lars Kunze

To safely navigate intricate real-world scenarios, autonomous vehicles must be able to adapt to diverse road conditions and anticipate future events. World model (WM) based reinforcement learning (RL) has emerged as a promising approach by…

机器人学 · 计算机科学 2024-07-29 Dechen Gao , Shuangyu Cai , Hanchu Zhou , Hang Wang , Iman Soltani , Junshan Zhang

World generation is a fundamental capability for applications like video games, simulation, and robotics. However, existing approaches face three main obstacles: controllability, scalability, and efficiency. End-to-end scene generation…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Han-Hung Lee , Cheng-Yu Yang , Yu-Lun Liu , Angel X. Chang

Autonomous GUI agents interact with environments by perceiving interfaces and executing actions. As a virtual sandbox, the GUI World model empowers agents with human-like foresight by enabling action-conditioned prediction. However,…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Yuhao Zheng , Li'an Zhong , Yi Wang , Rui Dai , Kaikui Liu , Xiangxiang Chu , Linyuan Lv , Philip Torr , Kevin Qinghong Lin

End-to-end autonomous driving has received increasing attention due to its potential to learn from large amounts of data. However, most existing methods are still open-loop and suffer from weak scalability, lack of high-order interactions,…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Wenzhao Zheng , Zetian Xia , Yuanhui Huang , Sicheng Zuo , Jie Zhou , Jiwen Lu

Human drivers rely on commonsense reasoning to navigate diverse and dynamic real-world scenarios. Existing end-to-end (E2E) autonomous driving (AD) models are typically optimized to mimic driving patterns observed in data, without capturing…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Yi Xu , Yuxin Hu , Zaiwei Zhang , Gregory P. Meyer , Siva Karthik Mustikovela , Siddhartha Srinivasa , Eric M. Wolff , Xin Huang

Representation learning seeks to expose certain aspects of observed data in a learned representation that's amenable to downstream tasks like classification. For instance, a good representation for 2D images might be one that describes only…

Typically, autonomous driving adopts a modular design, which divides the full stack into perception, prediction, planning and control parts. Though interpretable, such modular design tends to introduce a substantial amount of redundancy.…

计算机视觉与模式识别 · 计算机科学 2023-11-23 Fan Jia , Weixin Mao , Yingfei Liu , Yucheng Zhao , Yuqing Wen , Chi Zhang , Xiangyu Zhang , Tiancai Wang

Generative video models, a leading approach to world modeling, face fundamental limitations. They often violate physical and logical rules, lack interactivity, and operate as opaque black boxes ill-suited for building structured, queryable…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Felix O'Mahony , Roberto Cipolla , Ayush Tewari