English
Related papers

Related papers: Dissecting RGB-D Learning for Improved Multi-modal…

200 papers

Correspondence estimation is one of the most widely researched and yet only partially solved area of computer vision with many applications in tracking, mapping, recognition of objects and environment. In this paper, we propose a novel way…

Computer Vision and Pattern Recognition · Computer Science 2020-04-16 Umashankar Deekshith , Nishit Gajjar , Max Schwarz , Sven Behnke

Recently, learning-based approaches show promising results in navigation tasks. However, the poor generalization capability and the simulation-reality gap prevent a wide range of applications. We consider the problem of improving the…

Robotics · Computer Science 2023-09-26 Wenzhe Cai , Guangran Cheng , Lingyue Kong , Lu Dong , Changyin Sun

RGB-infrared person re-identification is a challenging task due to the intra-class variations and cross-modality discrepancy. Existing works mainly focus on learning modality-shared global representations by aligning image styles or feature…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Junhui Yin , Zhanyu Ma , Jiyang Xie , Shibo Nie , Kongming Liang , Jun Guo

State-of-the-art LiDAR-camera 3D object detectors usually focus on feature fusion. However, they neglect the factor of depth while designing the fusion strategy. In this work, we are the first to observe that different modalities play…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Mingqian Ji , Jian Yang , Shanshan Zhang

Recognizing objects and scenes are two challenging but essential tasks in image understanding. In particular, the use of RGB-D sensors in handling these tasks has emerged as an important area of focus for better visual understanding.…

Computer Vision and Pattern Recognition · Computer Science 2022-01-12 Ali Caglayan , Nevrez Imamoglu , Ahmet Burak Can , Ryosuke Nakamura

As robotic technologies advancing towards more complex multimodal interactions and manipulation tasks, the integration of advanced Vision-Language Models (VLMs) has become a key driver in the field. Despite progress with current methods,…

Robotics · Computer Science 2025-03-26 Sheng Wang

Multispectral object detection aims to leverage complementary information from visible (RGB) and infrared (IR) modalities to enable robust performance under diverse environmental conditions. Our key insight, derived from wavelet analysis…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Seongmin Hwang , Daeyoung Han , Moongu Jeon

RGB-D salient object detection (SOD) aims to detect the prominent regions by jointly modeling RGB and depth information. Most RGB-D SOD methods apply the same type of backbones and fusion modules to identically learn the multimodality and…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Kang Yi , Jing Xu , Xiao Jin , Fu Guo , Yan-Feng Wu

The growing availability of commodity RGB-D cameras has boosted the applications in the field of scene understanding. However, as a fundamental scene understanding task, surface normal estimation from RGB-D data lacks thorough…

Computer Vision and Pattern Recognition · Computer Science 2019-11-21 Jin Zeng , Yanfeng Tong , Yunmu Huang , Qiong Yan , Wenxiu Sun , Jing Chen , Yongtian Wang

Given the widespread adoption of depth-sensing acquisition devices, RGB-D videos and related data/media have gained considerable traction in various aspects of daily life. Consequently, conducting salient object detection (SOD) in RGB-D…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Ao Mou , Yukang Lu , Jiahao He , Dingyao Min , Keren Fu , Qijun Zhao

Many machine learning problems concern with discovering or associating common patterns in data of multiple views or modalities. Multi-view learning is of the methods to achieve such goals. Recent methods propose deep multi-view networks via…

Computer Vision and Pattern Recognition · Computer Science 2019-09-04 Kui Jia , Jiehong Lin , Mingkui Tan , Dacheng Tao

Multi-object tracking from RGB-D video sequences is a challenging problem due to the combination of changing viewpoints, motion, and occlusions over time. We observe that having the complete geometry of objects aids in their tracking, and…

Computer Vision and Pattern Recognition · Computer Science 2020-12-17 Norman Müller , Yu-Shiang Wong , Niloy J. Mitra , Angela Dai , Matthias Nießner

The convergence of cross-modal adversarial learning and physics-driven methods represents a cutting-edge direction for tackling challenges in complex multi-modal tasks and scientific computing. This review focuses on systematically…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Hana Satou , Alan Mitkiy

We present a method for temporally consistent motion segmentation from RGB-D videos assuming a piecewise rigid motion model. We formulate global energies over entire RGB-D sequences in terms of the segmentation of each frame into a number…

Computer Vision and Pattern Recognition · Computer Science 2016-08-17 Peter Bertholet , Alexandru-Eugen Ichim , Matthias Zwicker

Multimodal MRIs play a crucial role in clinical diagnosis and treatment. Feature disentanglement (FD)-based methods, aiming at learning superior feature representations for multimodal data analysis, have achieved significant success in…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Tianling Liu , Hongying Liu , Fanhua Shang , Lequan Yu , Tong Han , Liang Wan

Image segmentation is a vital task for providing human assistance and enhancing autonomy in our daily lives. In particular, RGB-D segmentation-leveraging both visual and depth cues-has attracted increasing attention as it promises richer…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Aecheon Jung , Soyun Choi , Junhong Min , Sungeun Hong

The integration of RGB and depth modalities significantly enhances the accuracy of segmenting complex indoor scenes, with depth data from RGB-D cameras playing a crucial role in this improvement. However, collecting an RGB-D dataset is more…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Xinhua Xu , Hong Liu , Jianbing Wu , Jinfu Liu

Recognizing objects from simultaneously sensed photometric (RGB) and depth channels is a fundamental yet practical problem in many machine vision applications such as robot grasping and autonomous driving. In this paper, we address this…

Computer Vision and Pattern Recognition · Computer Science 2018-12-26 Guanbin Li , Yukang Gan , Hejun Wu , Nong Xiao , Liang Lin

The focal point of egocentric video understanding is modelling hand-object interactions. Standard models -- CNNs, Vision Transformers, etc. -- which receive RGB frames as input perform well, however, their performance improves further by…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Gorjan Radevski , Dusan Grujicic , Matthew Blaschko , Marie-Francine Moens , Tinne Tuytelaars

Applying data-driven approaches to non-rigid 3D reconstruction has been difficult, which we believe can be attributed to the lack of a large-scale training corpus. Unfortunately, this method fails for important cases such as highly…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Aljaž Božič , Michael Zollhöfer , Christian Theobalt , Matthias Nießner
‹ Prev 1 8 9 10 Next ›