English
Related papers

Related papers: Dissecting RGB-D Learning for Improved Multi-modal…

200 papers

Semantic segmentation has made striking progress due to the success of deep convolutional neural networks. Considering the demands of autonomous driving, real-time semantic segmentation has become a research hotspot these years. However,…

Computer Vision and Pattern Recognition · Computer Science 2020-06-30 Lei Sun , Kailun Yang , Xinxin Hu , Weijian Hu , Kaiwei Wang

Imitation learning has emerged as a crucial ap proach for acquiring visuomotor skills from demonstrations, where designing effective observation encoders is essential for policy generalization. However, existing methods often struggle to…

Robotics · Computer Science 2025-12-01 Yikai Tang , Haoran Geng , Sheng Zang , Pieter Abbeel , Jitendra Malik

The extensive research leveraging RGB-D information has been exploited in salient object detection. However, salient visual cues appear in various scales and resolutions of RGB images due to semantic gaps at different feature levels.…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Ze-yu Liu , Jian-wei Liu , Xin Zuo , Ming-fei Hu

RGB-D saliency detection aims to fuse multi-modal cues to accurately localize salient regions. Existing works often adopt attention modules for feature modeling, with few methods explicitly leveraging fine-grained details to merge with…

Computer Vision and Pattern Recognition · Computer Science 2023-04-19 Zongwei Wu , Guillaume Allibert , Fabrice Meriaudeau , Chao Ma , Cédric Demonceaux

2D face recognition encounters challenges in unconstrained environments due to varying illumination, occlusion, and pose. Recent studies focus on RGB-D face recognition to improve robustness by incorporating depth information. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Zijian Chen , Mei Wang , Weihong Deng , Hongzhi Shi , Dongchao Wen , Yingjie Zhang , Xingchen Cui , Jian Zhao

Autonomous agents that rely purely on perception to make real-time control decisions require efficient and robust architectures. In this work, we demonstrate that augmenting RGB input with depth information significantly enhances our…

Robotics · Computer Science 2025-11-14 Mihaela-Larisa Clement , Mónika Farsang , Felix Resch , Mihai-Teodor Stanusoiu , Radu Grosu

Existing RGB-D salient object detection (SOD) models usually treat RGB and depth as independent information and design separate networks for feature extraction from each. Such schemes can easily be constrained by a limited amount of…

Computer Vision and Pattern Recognition · Computer Science 2021-04-19 Keren Fu , Deng-Ping Fan , Ge-Peng Ji , Qijun Zhao , Jianbing Shen , Ce Zhu

The integration of dual-modal features has been pivotal in advancing RGB-Depth (RGB-D) tracking. However, current trackers are less efficient and focus solely on single-level features, resulting in weaker robustness in fusion and slower…

Computer Vision and Pattern Recognition · Computer Science 2025-04-25 Boyue Xu , Yi Xu , Ruichao Hou , Jia Bei , Tongwei Ren , Gangshan Wu

The multi-modal salient object detection model based on RGB-D information has better robustness in the real world. However, it remains nontrivial to better adaptively balance effective multi-modal information in the feature fusion phase. In…

Computer Vision and Pattern Recognition · Computer Science 2022-02-09 Jinchao Zhu , Xiaoyu Zhang , Xian Fang , Feng Dong , Qiu Yu

We present MVD-Fusion: a method for single-view 3D inference via generative modeling of multi-view-consistent RGB-D images. While recent methods pursuing 3D inference advocate learning novel-view generative models, these generations are not…

Computer Vision and Pattern Recognition · Computer Science 2024-04-05 Hanzhe Hu , Zhizhuo Zhou , Varun Jampani , Shubham Tulsiani

How to perform effective information fusion of different modalities is a core factor in boosting the performance of RGBT tracking. This paper presents a novel deep fusion algorithm based on the representations from an end-to-end trained…

Computer Vision and Pattern Recognition · Computer Science 2019-08-12 Yabin Zhu , Chenglong Li , Bin Luo , Jin Tang , Xiao Wang

Moving Object Detection (MOD) is a critical vision task for successfully achieving safe autonomous driving. Despite plausible results of deep learning methods, most existing approaches are only frame-based and may fail to reach reasonable…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Zhuyun Zhou , Zongwei Wu , Rémi Boutteau , Fan Yang , Cédric Demonceaux , Dominique Ginhac

Temporal action detection aims to predict the time intervals and the classes of action instances in the video. Despite the promising performance, existing two-stream models exhibit slow inference speed due to their reliance on…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Pilhyeon Lee , Taeoh Kim , Minho Shim , Dongyoon Wee , Hyeran Byun

The high performance of RGB-D based road segmentation methods contrasts with their rare application in commercial autonomous driving, which is owing to two reasons: 1) the prior methods cannot achieve high inference speed and high accuracy…

Computer Vision and Pattern Recognition · Computer Science 2022-03-10 Yicong Chang , Feng Xue , Fei Sheng , Wenteng Liang , Anlong Ming

The goal of multi-modal learning is to use complimentary information on the relevant task provided by the multiple modalities to achieve reliable and robust performance. Recently, deep learning has led significant improvement in multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2018-11-05 Jaekyum Kim , Junho Koh , Yecheol Kim , Jaehyung Choi , Youngbae Hwang , Jun Won Choi

Vision-based autonomous driving requires reliable and efficient object detection. This work proposes a DiffusionDet-based framework that exploits data fusion from the monocular camera and depth sensor to provide the RGB and depth (RGB-D)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Eliraz Orfaig , Inna Stainvas , Igal Bilik

Depth information available from an RGB-D camera can be useful in segmenting salient objects when figure/ground cues from RGB channels are weak. This has motivated the development of several RGB-D saliency datasets and algorithms that use…

Computer Vision and Pattern Recognition · Computer Science 2020-10-27 Yue Wang , Yuke Li , James H. Elder , Huchuan Lu , Runmin Wu , Lu Zhang

RGB-guided depth completion aims at predicting dense depth maps from sparse depth measurements and corresponding RGB images, where how to effectively and efficiently exploit the multi-modal information is a key issue. Guided dynamic…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Yufei Wang , Yuxin Mao , Qi Liu , Yuchao Dai

Human motion recognition is one of the most important branches of human-centered research activities. In recent years, motion recognition based on RGB-D data has attracted much attention. Along with the development in artificial…

Computer Vision and Pattern Recognition · Computer Science 2018-04-26 Pichao Wang , Wanqing Li , Philip Ogunbona , Jun Wan , Sergio Escalera

We present RGB-D-Fusion, a multi-modal conditional denoising diffusion probabilistic model to generate high resolution depth maps from low-resolution monocular RGB images of humanoid subjects. RGB-D-Fusion first generates a low-resolution…

Computer Vision and Pattern Recognition · Computer Science 2023-09-25 Sascha Kirch , Valeria Olyunina , Jan Ondřej , Rafael Pagés , Sergio Martin , Clara Pérez-Molina
‹ Prev 1 3 4 5 6 7 10 Next ›