English
Related papers

Related papers: OccLoff: Learning Optimized Feature Fusion for 3D …

200 papers

With the advent of deep neural networks, learning-based approaches for 3D reconstruction have gained popularity. However, unlike for images, in 3D there is no canonical representation which is both computationally and memory efficient yet…

Computer Vision and Pattern Recognition · Computer Science 2019-05-01 Lars Mescheder , Michael Oechsle , Michael Niemeyer , Sebastian Nowozin , Andreas Geiger

Learning-based 3D Scanning plays a crucial role in enabling efficient and accurate scanning of target objects. However, recent reinforcement learning-based methods often require large-scale training data and still struggle to generalize to…

Robotics · Computer Science 2026-03-12 Itsuki Hirako , Ryo Hakoda , Yubin Liu , Matthew Hwang , Yoshihiro Sato , Takeshi Oishi

In this paper, we present a learning based approach to depth fusion, i.e., dense 3D reconstruction from multiple depth images. The most common approach to depth fusion is based on averaging truncated signed distance functions, which was…

Computer Vision and Pattern Recognition · Computer Science 2017-11-02 Gernot Riegler , Ali Osman Ulusoy , Horst Bischof , Andreas Geiger

This technical report summarizes the winning solution for the 3D Occupancy Prediction Challenge, which is held in conjunction with the CVPR 2023 Workshop on End-to-End Autonomous Driving and CVPR 23 Workshop on Vision-Centric Autonomous…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Zhiqi Li , Zhiding Yu , David Austin , Mingsheng Fang , Shiyi Lan , Jan Kautz , Jose M. Alvarez

Developing effective multimodal data fusion strategies has become increasingly essential for improving the predictive power of statistical machine learning methods across a wide range of applications, from autonomous driving to medical…

Machine Learning · Computer Science 2025-07-29 Ziyi Liang , Annie Qu , Babak Shahbaba

We propose a late-to-early recurrent feature fusion scheme for 3D object detection using temporal LiDAR point clouds. Our main motivation is fusing object-aware latent embeddings into the early stages of a 3D object detector. This feature…

Computer Vision and Pattern Recognition · Computer Science 2023-10-02 Tong He , Pei Sun , Zhaoqi Leng , Chenxi Liu , Dragomir Anguelov , Mingxing Tan

Online 3D occupancy prediction provides a comprehensive spatial understanding of embodied environments. While the innovative EmbodiedOcc framework utilizes 3D semantic Gaussians for progressive indoor occupancy prediction, it overlooks the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Hao Wang , Xiaobao Wei , Xiaoan Zhang , Jianing Li , Chengyu Bai , Ying Li , Ming Lu , Wenzhao Zheng , Shanghang Zhang

The task of motion prediction is pivotal for autonomous driving systems, providing crucial data to choose a vehicle behavior strategy within its surroundings. Existing motion prediction techniques primarily focus on predicting the future…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Youshaa Murhij , Dmitry Yudin

While 3D object bounding box (bbox) representation has been widely used in autonomous driving perception, it lacks the ability to capture the precise details of an object's intrinsic geometry. Recently, occupancy has emerged as a promising…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Chaoda Zheng , Feng Wang , Naiyan Wang , Shuguang Cui , Zhen Li

Achieving highly accurate and real-time 3D occupancy prediction from cameras is a critical requirement for the safe and practical deployment of autonomous vehicles. While this shift to sparse 3D representations solves the encoding…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Suzeyu Chen , Leheng Li , Ying-Cong Chen

Understanding and forecasting the scene evolutions deeply affect the exploration and decision of embodied agents. While traditional methods simulate scene evolutions through trajectory prediction of potential instances, current works use…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Zhang Zhang , Qiang Zhang , Wei Cui , Shuai Shi , Yijie Guo , Gang Han , Wen Zhao , Jingkai Sun , Jiahang Cao , Jiaxu Wang , Hao Cheng , Xiaozhu Ju , Zhengping Che , Renjing Xu , Jian Tang

In autonomous driving, recent research has increasingly focused on collaborative perception based on deep learning to overcome the limitations of individual perception systems. Although these methods achieve high accuracy, they rely on high…

Robotics · Computer Science 2025-07-04 Maryem Fadili , Mohamed Anis Ghaoui , Louis Lecrosnier , Steve Pechberti , Redouane Khemmar

Recognizing 3D part instances from a 3D point cloud is crucial for 3D structure and scene understanding. Several learning-based approaches use semantic segmentation and instance center prediction as training tasks and fail to further…

Computer Vision and Pattern Recognition · Computer Science 2022-08-10 Chunyu Sun , Xin Tong , Yang Liu

3D instance segmentation, with a variety of applications in robotics and augmented reality, is in large demands these days. Unlike 2D images that are projective observations of the environment, 3D models provide metric reconstruction of the…

Computer Vision and Pattern Recognition · Computer Science 2020-04-29 Lei Han , Tian Zheng , Lan Xu , Lu Fang

We present SOccDPT, a memory-efficient approach for 3D semantic occupancy prediction from monocular image input using dense prediction transformers. To address the limitations of existing methods trained on structured traffic datasets, we…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Aditya Nalgunda Ganesh

Occupancy is crucial for autonomous driving, providing essential geometric priors for perception and planning. However, existing methods predominantly rely on LiDAR-based occupancy annotations, which limits scalability and prevents…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Baijun Ye , Minghui Qin , Saining Zhang , Moonjun Gong , Shaoting Zhu , Zebang Shen , Luan Zhang , Lu Zhang , Hao Zhao , Hang Zhao

State-of-the-art LiDAR-camera 3D object detectors usually focus on feature fusion. However, they neglect the factor of depth while designing the fusion strategy. In this work, we are the first to observe that different modalities play…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Mingqian Ji , Jian Yang , Shanshan Zhang

Multi-modal systems enhance performance in autonomous driving but face inefficiencies due to indiscriminate processing within each modality. Additionally, the independent feature learning of each modality lacks interaction, which results in…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Guoliang You , Xiaomeng Chu , Yifan Duan , Xingchen Li , Sha Zhang , Jianmin Ji , Yanyong Zhang

In the field of 3D Human Pose Estimation from monocular videos, the presence of diverse occlusion types presents a formidable challenge. Prior research has made progress by harnessing spatial and temporal cues to infer 3D poses from 2D…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Mehwish Ghafoor , Arif Mahmood , Muhammad Bilal

A real-time semantic 3D occupancy mapping framework is proposed in this paper. The mapping framework is based on the Bayesian kernel inference strategy from the literature. Two novel free space representations are proposed to efficiently…

Robotics · Computer Science 2021-07-08 Yuanxin Zhong , Huei Peng