English
Related papers

Related papers: OC-SOP: Enhancing Vision-Based 3D Semantic Occupan…

200 papers

Scene completion and forecasting are two popular perception problems in research for mobile agents like autonomous vehicles. Existing approaches treat the two problems in isolation, resulting in a separate perception of the two aspects. In…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Xinhao Liu , Moonjun Gong , Qi Fang , Haoyu Xie , Yiming Li , Hang Zhao , Chen Feng

Accurate perception of dynamic traffic scenes is crucial for high-level autonomous driving systems, requiring robust object motion estimation and instance segmentation. However, traditional methods often treat them as separate tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Yinqi Chen , Meiying Zhang , Qi Hao , Guang Zhou

Obtaining high-quality 3D semantic occupancy from raw sensor data remains an essential yet challenging task, often requiring extensive manual labeling. In this work, we propose AutoOcc, a vision-centric automated pipeline for open-ended…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Xiaoyu Zhou , Jingqi Wang , Yongtao Wang , Yufei Wei , Nan Dong , Ming-Hsuan Yang

Image matching is a fundamental and critical task in various visual applications, such as Simultaneous Localization and Mapping (SLAM) and image retrieval, which require accurate pose estimation. However, most existing methods ignore the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Miao Fan , Mingrui Chen , Chen Hu , Shuchang Zhou

Object-centric representation learning aims to decompose visual scenes into fixed-size vectors called "slots" or "object files", where each slot captures a distinct object. Current state-of-the-art object-centric models have shown…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Aniket Didolkar , Andrii Zadaianchuk , Rabiul Awal , Maximilian Seitzer , Efstratios Gavves , Aishwarya Agrawal

Recent advances in Multi-modal Large Language Models (MLLMs) have showcased remarkable capabilities in vision-language understanding. However, enabling robust video spatial reasoning-the ability to comprehend object locations, orientations,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Haoran Tang , Meng Cao , Ruyang Liu , Xiaoxi Liang , Linglong Li , Ge Li , Xiaodan Liang

Autonomous driving systems require a quick and robust perception of the nearby environment to carry out their routines effectively. With the aim to avoid collisions and drive safely, autonomous driving systems rely heavily on object…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Abdul Hannan Khan , Syed Tahseen Raza Rizvi , Dheeraj Varma Chittari Macharavtu , Andreas Dengel

3D reconstruction has been widely used in autonomous navigation fields of mobile robotics. However, the former research can only provide the basic geometry structure without the capability of open-world scene understanding, limiting…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Haochen Jiang , Yueming Xu , Yihan Zeng , Hang Xu , Wei Zhang , Jianfeng Feng , Li Zhang

Understanding world dynamics is crucial for planning in autonomous driving. Recent methods attempt to achieve this by learning a 3D occupancy world model that forecasts future surrounding scenes based on current observation. However, 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-02-12 Xiang Li , Pengfei Li , Yupeng Zheng , Wei Sun , Yan Wang , Yilun Chen

Given two consecutive frames from a pair of stereo cameras, 3D scene flow methods simultaneously estimate the 3D geometry and motion of the observed scene. Many existing approaches use superpixels for regularization, but may predict…

Computer Vision and Pattern Recognition · Computer Science 2017-10-09 Zhile Ren , Deqing Sun , Jan Kautz , Erik B. Sudderth

Learning an egocentric action recognition model from video data is challenging due to distractors (e.g., irrelevant objects) in the background. Further integrating object information into an action model is hence beneficial. Existing…

Computer Vision and Pattern Recognition · Computer Science 2022-05-04 Victor Escorcia , Ricardo Guerrero , Xiatian Zhu , Brais Martinez

Multi-sensor fusion significantly enhances the accuracy and robustness of 3D semantic occupancy prediction, which is crucial for autonomous driving and robotics. However, most existing approaches depend on high-resolution images and complex…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Zhen Yang , Yanpeng Dong , Jiayu Wang , Heng Wang , Lichao Ma , Zijian Cui , Qi Liu , Haoran Pei , Kexin Zhang , Chao Zhang

Visual simultaneous localization and mapping (SLAM) systems face challenges in detecting loop closure under the circumstance of large viewpoint changes. In this paper, we present an object-based loop closure detection method based on the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-17 Xingwu Ji , Peilin Liu , Haochen Niu , Xiang Chen , Rendong Ying , Fei Wen

Developing 3D semantic occupancy prediction models often relies on dense 3D annotations for supervised learning, a process that is both labor and resource-intensive, underscoring the need for label-efficient or even label-free approaches.…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Samuel Sze , Daniele De Martini , Lars Kunze

The vision-based perception for autonomous driving has undergone a transformation from the bird-eye-view (BEV) representations to the 3D semantic occupancy. Compared with the BEV planes, the 3D semantic occupancy further provides structural…

Computer Vision and Pattern Recognition · Computer Science 2023-04-12 Yunpeng Zhang , Zheng Zhu , Dalong Du

Open-Vocabulary Object Detection (OVOD) aims to generalize object recognition to novel categories, while Weakly Supervised OVOD (WS-OVOD) extends this by combining box-level annotations with image-level labels. Despite recent progress, two…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Jiaying Zhou , Qingchao Chen

Semantic segmentation is a powerful method to facilitate visual scene understanding. Each pixel is assigned a label according to a pre-defined list of object classes and semantic entities. This becomes very useful as a means to summarize…

Computer Vision and Pattern Recognition · Computer Science 2018-11-21 Marc Bosch , Gordon A. Christie , Christopher M. Gifford

Recent remote sensing tech advancements drive imagery growth, making oriented object detection rapid development, yet hindered by labor-intensive annotation for high-density scenes. Oriented object detection with point supervision offers a…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Xinyuan Liu , Hang Xu , Yike Ma , Yucheng Zhang , Feng Dai

Manually specifying features that capture the diversity in traffic environments is impractical. Consequently, learning-based agents cannot realize their full potential as neural motion planners for autonomous vehicles. Instead, this work…

Machine Learning · Computer Science 2023-03-09 Eivind Meyer , Lars Frederik Peiss , Matthias Althoff

We present SOccDPT, a memory-efficient approach for 3D semantic occupancy prediction from monocular image input using dense prediction transformers. To address the limitations of existing methods trained on structured traffic datasets, we…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Aditya Nalgunda Ganesh