English
Related papers

Related papers: Robust Double-Encoder Network for RGB-D Panoptic S…

200 papers

We show that it is possible to learn semantic segmentation from very limited amounts of manual annotations, by enforcing geometric 3D constraints between multiple views. More exactly, image locations corresponding to the same physical 3D…

Computer Vision and Pattern Recognition · Computer Science 2019-01-10 Sinisa Stekovic , Friedrich Fraundorfer , Vincent Lepetit

Indoor scene semantic parsing from RGB images is very challenging due to occlusions, object distortion, and viewpoint variations. Going beyond prior works that leverage geometry information, typically paired depth maps, we present a new…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Zhengzhe Liu , Xiaojuan Qi , Chi-Wing Fu

Raw depth images captured in indoor scenarios frequently exhibit extensive missing values due to the inherent limitations of the sensors and environments. For example, transparent materials frequently elude detection by depth sensors;…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Haowen Wang , Zhengping Che , Yufan Yang , Mingyuan Wang , Zhiyuan Xu , Xiuquan Qiao , Mengshi Qi , Feifei Feng , Jian Tang

Glass surfaces are becoming increasingly ubiquitous as modern buildings tend to use a lot of glass panels. This, however, poses substantial challenges to the operations of autonomous systems such as robots, self-driving cars, and drones, as…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Jiaying Lin , Yuen-Hei Yeung , Shuquan Ye , Rynson W. H. Lau

Gesture recognition is getting more and more popular due to various application possibilities in human-machine interaction. Existing multi-modal gesture recognition systems take multi-modal data as input to improve accuracy, but such…

Computer Vision and Pattern Recognition · Computer Science 2021-11-01 Dinghao Fan , Hengjie Lu , Shugong Xu , Shan Cao

In this paper, we propose a new progressive pre-training method for image understanding tasks which leverages RGB-D datasets. The method utilizes Multi-Modal Contrastive Masked Autoencoder and Denoising techniques. Our proposed approach…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Muhammad Abdullah Jamal , Omid Mohareri

We propose a real-time general purpose semantic segmentation architecture, RGPNet, which achieves significant performance gain in complex environments. RGPNet consists of a light-weight asymmetric encoder-decoder and an adaptor. The adaptor…

Computer Vision and Pattern Recognition · Computer Science 2020-12-16 Elahe Arani , Shabbir Marzban , Andrei Pata , Bahram Zonooz

Semantic segmentation networks are usually pre-trained once and not updated during deployment. As a consequence, misclassifications commonly occur if the distribution of the training data deviates from the one encountered during the robot's…

Robotics · Computer Science 2023-02-15 Jonas Frey , Hermann Blum , Francesco Milano , Roland Siegwart , Cesar Cadena

Reconstructing three-dimensional (3D) scenes with semantic understanding is vital in many robotic applications. Robots need to identify which objects, along with their positions and shapes, to manipulate them precisely with given tasks.…

Robotics · Computer Science 2024-12-17 Khang Nguyen , Tuan Dang , Manfred Huber

We present a deep reinforcement learning method of progressive view inpainting for colored semantic point cloud scene completion under volume guidance, achieving high-quality scene reconstruction from only a single RGB-D image with severe…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Zhaoxuan Zhang , Xiaoguang Han , Bo Dong , Tong Li , Baocai Yin , Xin Yang

Image segmentation for video analysis plays an essential role in different research fields such as smart city, healthcare, computer vision and geoscience, and remote sensing applications. In this regard, a significant effort has been…

Computer Vision and Pattern Recognition · Computer Science 2021-11-22 Omar Elharrouss , Somaya Al-Maadeed , Nandhini Subramanian , Najmath Ottakath , Noor Almaadeed , Yassine Himeur

The multi-modal salient object detection model based on RGB-D information has better robustness in the real world. However, it remains nontrivial to better adaptively balance effective multi-modal information in the feature fusion phase. In…

Computer Vision and Pattern Recognition · Computer Science 2022-02-09 Jinchao Zhu , Xiaoyu Zhang , Xian Fang , Feng Dong , Qiu Yu

Robot navigation in mapless environment is one of the essential problems and challenges in mobile robots. Deep reinforcement learning is a promising technique to tackle the task of mapless navigation. Since reinforcement learning requires a…

Robotics · Computer Science 2019-04-23 Liulong Ma , Yanjie Liu* , Jiao Chen

This paper addresses the issue on how to more effectively coordinate the depth with RGB aiming at boosting the performance of RGB-D object detection. Particularly, we investigate two primary ideas under the CNN model: property derivation…

Computer Vision and Pattern Recognition · Computer Science 2016-05-10 Saihui Hou , Zilei Wang , Feng Wu

Existing RGB-D salient object detection (SOD) approaches concentrate on the cross-modal fusion between the RGB stream and the depth stream. They do not deeply explore the effect of the depth map itself. In this work, we design a single…

Computer Vision and Pattern Recognition · Computer Science 2020-07-16 Xiaoqi Zhao , Lihe Zhang , Youwei Pang , Huchuan Lu , Lei Zhang

Depth estimation is crucial for intelligent systems, enabling applications from autonomous navigation to augmented reality. While traditional stereo and active depth sensors have limitations in cost, power, and robustness, dual-pixel (DP)…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Kunal Swami , Debtanu Gupta , Amrit Kumar Muduli , Chirag Jaiswal , Pankaj Kumar Bajpai

Multimodal deep sensor fusion has the potential to enable autonomous vehicles to visually understand their surrounding environments in all weather conditions. However, existing deep sensor fusion methods usually employ convoluted…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Sri Aditya Deevi , Connor Lee , Lu Gan , Sushruth Nagesh , Gaurav Pandey , Soon-Jo Chung

Recovering the radiometric properties of a scene (i.e., the reflectance, illumination, and geometry) is a long-sought ability of computer vision that can provide invaluable information for a wide range of applications. Deciphering the…

Computer Vision and Pattern Recognition · Computer Science 2016-04-06 Stephen Lombardi , Ko Nishino

Understanding the scene in which an autonomous robot operates is critical for its competent functioning. Such scene comprehension necessitates recognizing instances of traffic participants along with general scene semantics which can be…

Computer Vision and Pattern Recognition · Computer Science 2021-11-05 Rohit Mohan , Abhinav Valada

We present 3DMV, a novel method for 3D semantic scene segmentation of RGB-D scans in indoor environments using a joint 3D-multi-view prediction network. In contrast to existing methods that either use geometry or RGB data as input for this…

Computer Vision and Pattern Recognition · Computer Science 2018-03-29 Angela Dai , Matthias Nießner