English
Related papers

Related papers: GaussianFormer3D: Multi-Modal Gaussian-based Seman…

200 papers

3D open-vocabulary scene understanding, crucial for advancing augmented reality and robotic applications, involves interpreting and locating specific regions within a 3D space as directed by natural language instructions. To this end, we…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Yansong Qu , Shaohui Dai , Xinyang Li , Jianghang Lin , Liujuan Cao , Shengchuan Zhang , Rongrong Ji

We present 3D Spatial MultiModal Memory (M3), a multimodal memory system designed to retain information about medium-sized static scenes through video sources for visual perception. By integrating 3D Gaussian Splatting techniques with…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Xueyan Zou , Yuchen Song , Ri-Zhao Qiu , Xuanbin Peng , Jianglong Ye , Sifei Liu , Xiaolong Wang

Monocular 3D detection is a challenging task due to the lack of accurate 3D information. Existing approaches typically rely on geometry constraints and dense depth estimates to facilitate the learning, but often fail to fully exploit the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Liang Peng , Junkai Xu , Haoran Cheng , Zheng Yang , Xiaopei Wu , Wei Qian , Wenxiao Wang , Boxi Wu , Deng Cai

Recent developments in 3D reconstruction and neural rendering have significantly propelled the capabilities of photo-realistic 3D scene rendering across various academic and industrial fields. The 3D Gaussian Splatting technique, alongside…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Zexu Huang , Min Xu , Stuart Perry

Recent progress in pre-trained diffusion models and 3D generation have spurred interest in 4D content creation. However, achieving high-fidelity 4D generation with spatial-temporal consistency remains a challenge. In this work, we propose…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Yifei Zeng , Yanqin Jiang , Siyu Zhu , Yuanxun Lu , Youtian Lin , Hao Zhu , Weiming Hu , Xun Cao , Yao Yao

Visual navigation for autonomous agents is a core task in the fields of computer vision and robotics. Learning-based methods, such as deep reinforcement learning, have the potential to outperform the classical solutions developed for this…

Computer Vision and Pattern Recognition · Computer Science 2021-03-23 Zachary Seymour , Kowshik Thopalli , Niluthpol Mithun , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

While multi-modal 3D semantic occupancy prediction typically enhances robustness by fusing camera and LiDAR inputs, its effectiveness is fundamentally constrained by environmental variability. Specifically, camera sensors suffer from severe…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 A. Enes Doruk , Abdelaziz Hussein , Hasan F. Ates

Recent advances in Gaussian Splatting based 3D scene representation have shown two major trends: semantics-oriented approaches that focus on high-level understanding but lack explicit 3D geometry modeling, and structure-oriented approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Yuhang Ming , Chenxin Fang , Xingyuan Yu , Fan Zhang , Weichen Dai , Wanzeng Kong , Guofeng Zhang

3D Gaussian Splatting, known for enabling high-quality static scene reconstruction with fast rendering, is increasingly being applied to multi-view dynamic scene reconstruction. A common strategy involves learning a deformation field to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Han Jiao , Jiakai Sun , Yexing Xu , Lei Zhao , Wei Xing , Huaizhong Lin

Mobility-on-demand (MoD) systems have recently emerged as a promising paradigm of one-way vehicle sharing for sustainable personal urban mobility in densely populated cities. In this paper, we enhance the capability of a MoD system by…

Robotics · Computer Science 2013-06-07 Jie Chen , Kian Hsiang Low , Colin Keng-Yan Tan

The task of 3D semantic scene completion using monocular cameras is gaining significant attention in the field of autonomous driving. This task aims to predict the occupancy status and semantic labels of each voxel in a 3D scene from…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Jiawei Yao , Jusheng Zhang , Xiaochao Pan , Tong Wu , Canran Xiao

Capturing and reconstructing high-speed dynamic 3D scenes has numerous applications in computer graphics, vision, and interdisciplinary fields such as robotics, aerodynamics, and evolutionary biology. However, achieving this using a single…

Computer Vision and Pattern Recognition · Computer Science 2025-02-10 Zihao Zou , Ziyuan Qu , Xi Peng , Vivek Boominathan , Adithya Pediredla , Praneeth Chakravarthula

While 3D object bounding box (bbox) representation has been widely used in autonomous driving perception, it lacks the ability to capture the precise details of an object's intrinsic geometry. Recently, occupancy has emerged as a promising…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Chaoda Zheng , Feng Wang , Naiyan Wang , Shuguang Cui , Zhen Li

Dynamic scene reconstruction is a long-term challenge in the field of 3D vision. Recently, the emergence of 3D Gaussian Splatting has provided new insights into this problem. Although subsequent efforts rapidly extend static 3D Gaussian to…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Ruijie Zhu , Yanzhe Liang , Hanzhi Chang , Jiacheng Deng , Jiahao Lu , Wenfei Yang , Tianzhu Zhang , Yongdong Zhang

Recent advancements in 3D scene understanding have made significant strides in enabling interaction with scenes using open-vocabulary queries, particularly for VR/AR and robotic applications. Nevertheless, existing methods are hindered by…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Dianyi Yang , Xihan Wang , Yu Gao , Shiyang Liu , Bohan Ren , Yufeng Yue , Yi Yang

The performance of multi-modal 3D occupancy prediction is limited by ineffective fusion, mainly due to geometry-semantics mismatch from fixed fusion strategies and surface detail loss caused by sparse, noisy annotations. The mismatch stems…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Luyao Lei , Shuo Xu , Yifan Bai , Xing Wei

3D semantic occupancy prediction is an essential part of autonomous driving, focusing on capturing the geometric details of scenes. Off-road environments are rich in geometric information, therefore it is suitable for 3D semantic occupancy…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Heng Zhai , Jilin Mei , Chen Min , Liang Chen , Fangzhou Zhao , Yu Hu

Simultaneous localization and mapping is essential for position tracking and scene understanding. 3D Gaussian-based map representations enable photorealistic reconstruction and real-time rendering of scenes using multiple posed cameras. We…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Lisong C. Sun , Neel P. Bhatt , Jonathan C. Liu , Zhiwen Fan , Zhangyang Wang , Todd E. Humphreys , Ufuk Topcu

Visual grounding aims to identify objects or regions in a scene based on natural language descriptions, essential for spatially aware perception in autonomous driving. However, existing visual grounding tasks typically depend on bounding…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Zhan Shi , Song Wang , Junbo Chen , Jianke Zhu

Camera-based occupancy prediction is a mainstream approach for 3D perception in autonomous driving, aiming to infer complete 3D scene geometry and semantics from 2D images. Almost existing methods focus on improving performance through…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Rongtao Xu , Jinzhou Lin , Jialei Zhou , Jiahua Dong , Changwei Wang , Ruisheng Wang , Li Guo , Shibiao Xu , Xiaodan Liang
‹ Prev 1 4 5 6 7 8 10 Next ›