English
Related papers

Related papers: GaussianWorld: Gaussian World Model for Streaming …

200 papers

Open-vocabulary 3D occupancy is vital for embodied agents, which need to understand complex indoor environments where semantic categories are abundant and evolve beyond fixed taxonomies. While recent work has explored open-vocabulary…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Changqing Zhou , Yueru Luo , Han Zhang , Zeyu Jiang , Changhao Chen

Recently, world models have made significant progress in enhancing end-to-end driving systems through both future situation forecasting and improved scene understanding. However, existing driving world models are typically built upon dense…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Ruoyu Wang , Jingke Wang , Yukai Ma , Yuehao Huang , Shuangming Lei , Guanglin Xu , Aixue Ye , Yong Liu

Understanding dynamic 3D environments is essential for safe autonomous driving, particularly when reasoning about human-centric, nonrigid agents. However, existing weakly supervised occupancy prediction frameworks predominantly assume…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Yang Gao , Wuyang Li , Po-Chien Luan , Alexandre Alahi

When exploring new areas, robotic systems generally exclusively plan and execute controls over geometry that has been directly measured. When entering space that was previously obstructed from view such as turning corners in hallways or…

Robotics · Computer Science 2024-03-19 Alec Reed , Brendan Crowe , Doncey Albin , Lorin Achey , Bradley Hayes , Christoffer Heckman

In this paper, we explore a novel point representation for 3D occupancy prediction from multi-view images, which is named Occupancy as Set of Points. Existing camera-based methods tend to exploit dense volume-based representation to predict…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Yiang Shi , Tianheng Cheng , Qian Zhang , Wenyu Liu , Xinggang Wang

Open-vocabulary scene understanding is crucial for robotic applications, enabling robots to comprehend complex 3D environmental contexts and supporting various downstream tasks such as navigation and manipulation. However, existing methods…

Autonomous driving requires a persistent understanding of 3D scenes that is robust to temporal disturbances and accounts for potential future actions. We introduce a new concept of 4D Occupancy Spatio-Temporal Persistence (OccSTeP), which…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Yu Zheng , Jie Hu , Kailun Yang , Jiaming Zhang

We propose DOME, a diffusion-based world model that predicts future occupancy frames based on past occupancy observations. The ability of this world model to capture the evolution of the environment is crucial for planning in autonomous…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Songen Gu , Wei Yin , Bu Jin , Xiaoyang Guo , Junming Wang , Haodong Li , Qian Zhang , Xiaoxiao Long

Instance-level change detection in 3D scenes presents significant challenges, particularly in uncontrolled environments lacking labeled image pairs, consistent camera poses, or uniform lighting conditions. This paper addresses these…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Binbin Jiang , Rui Huang , Qingyi Zhao , Yuxiang Zhang

Generating a coherent 3D scene representation from multi-view images is a fundamental yet challenging task. Existing methods often struggle with multi-view fusion, leading to fragmented 3D representations and sub-optimal performance. To…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Junho Kim , Seongwon Lee

Real-time, high-fidelity reconstruction of dynamic driving scenes is challenged by complex dynamics and sparse views, with prior methods struggling to balance quality and efficiency. We propose DrivingScene, an online, feed-forward…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Qirui Hou , Wenzhang Sun , Chang Zeng , Chunfeng Wang , Hao Li , Jianxun Cui

Efficient neural representations for dynamic video scenes are critical for applications ranging from video compression to interactive simulations. Yet, existing methods often face challenges related to high memory usage, lengthy training…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Andrew Bond , Jui-Hsien Wang , Long Mai , Erkut Erdem , Aykut Erdem

Human driver can easily describe the complex traffic scene by visual system. Such an ability of precise perception is essential for driver's planning. To achieve this, a geometry-aware representation that quantizes the physical 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2023-06-27 Chonghao Sima , Wenwen Tong , Tai Wang , Li Chen , Silei Wu , Hanming Deng , Yi Gu , Lewei Lu , Ping Luo , Dahua Lin , Hongyang Li

Open-vocabulary 3D scene understanding presents a significant challenge in computer vision, with wide-ranging applications in embodied agents and augmented reality systems. Existing methods adopt neurel rendering methods as 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Jun Guo , Xiaojian Ma , Yue Fan , Huaping Liu , Qing Li

Robotic perception requires the modeling of both 3D geometry and semantics. Existing methods typically focus on estimating 3D bounding boxes, neglecting finer geometric details and struggling to handle general, out-of-vocabulary objects. 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-12-14 Xiaoyu Tian , Tao Jiang , Longfei Yun , Yucheng Mao , Huitong Yang , Yue Wang , Yilun Wang , Hang Zhao

3D occupancy prediction has recently emerged as a new paradigm for holistic 3D scene understanding and provides valuable information for downstream planning in autonomous driving. Most existing methods, however, are computationally…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Yunxiao Shi , Hong Cai , Amin Ansari , Fatih Porikli

Occupancy prediction plays a pivotal role in autonomous driving. Previous methods typically construct dense 3D volumes, neglecting the inherent sparsity of the scene and suffering from high computational costs. To bridge the gap, we…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Haisong Liu , Yang Chen , Haiguang Wang , Zetong Yang , Tianyu Li , Jia Zeng , Li Chen , Hongyang Li , Limin Wang

End-to-end autonomous driving systems increasingly rely on vision-centric world models to understand and predict their environment. However, a common ineffectiveness in these models is the full reconstruction of future scenes, which expends…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Jianbiao Mei , Yu Yang , Xuemeng Yang , Licheng Wen , Jiajun Lv , Botian Shi , Yong Liu

Estimating 3D occupancy and motion at the vehicle's surroundings is essential for autonomous driving, enabling situational awareness in dynamic environments. Existing approaches jointly learn geometry and motion but rely on expensive 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Xavier Timoneda , Markus Herb , Fabian Duerr , Daniel Goehring

Accurately perceiving dynamic environments is a fundamental task for autonomous driving and robotic systems. Existing methods inadequately utilize temporal information, relying mainly on local temporal interactions between adjacent frames…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Tianhao Li , Yang Li , Mengtian Li , Yisheng Deng , Weifeng Ge