English
Related papers

Related papers: Efficient Bi-manipulation using RGBD Multi-model F…

200 papers

Reliable UAV object detection requires robustness to illumination changes, motion blur, and scene dynamics that suppress RGB cues. Thermal long-wave infrared (LWIR) sensing preserves contrast in low light, and event cameras retain…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Craig Iaboni , Pramod Abichandani

Multimodal object detection improves robustness in chal- lenging conditions by leveraging complementary cues from multiple sensor modalities. We introduce Filtered Multi- Modal Cross Attention Fusion (FMCAF), a preprocess- ing architecture…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Jad Berjawi , Yoann Dupas , Christophe C'erin

Existing RGB-Event detection methods process the low-information regions of both modalities (background in images and non-event regions in event data) uniformly during feature extraction and fusion, resulting in high computational costs and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Nan Yang , Yang Wang , Zhanwen Liu , Yuchao Dai , Yang Liu , Xiangmo Zhao

RGB-Thermal object tracking attempt to locate target object using complementary visual and thermal infrared data. Existing RGB-T trackers fuse different modalities by robust feature representation learning or adaptive modal weighting.…

Computer Vision and Pattern Recognition · Computer Science 2019-08-14 Rui Yang , Yabin Zhu , Xiao Wang , Chenglong Li , Jin Tang

Existing deep learning-based image inpainting methods typically rely on convolutional networks with RGB images to reconstruct images. However, relying exclusively on RGB images may neglect important depth information, which plays a critical…

Image and Video Processing · Electrical Eng. & Systems 2025-05-09 Jin Hyun Park , Harine Choi , Praewa Pitiphat

Autonomous agents that rely purely on perception to make real-time control decisions require efficient and robust architectures. In this work, we demonstrate that augmenting RGB input with depth information significantly enhances our…

Robotics · Computer Science 2025-11-14 Mihaela-Larisa Clement , Mónika Farsang , Felix Resch , Mihai-Teodor Stanusoiu , Radu Grosu

RGB-D saliency detection integrates information from both RGB images and depth maps to improve prediction of salient regions under challenging conditions. The key to RGB-D saliency detection is to fully mine and fuse information at multiple…

Computer Vision and Pattern Recognition · Computer Science 2021-12-02 Yue Wang , Xu Jia , Lu Zhang , Yuke Li , James Elder , Huchuan Lu

This paper presents an investigation into the estimation of optical and scene flow using RGBD information in scenarios where the RGB modality is affected by noise or captured in dark environments. Existing methods typically rely solely on…

Computer Vision and Pattern Recognition · Computer Science 2023-07-31 Youjie Zhou , Guofeng Mei , Yiming Wang , Fabio Poiesi , Yi Wan

Most existing RGB-D salient object detection (SOD) methods focus on the foreground region when utilizing the depth images. However, the background also provides important information in traditional SOD methods for promising performance. To…

Computer Vision and Pattern Recognition · Computer Science 2021-02-24 Zhao Zhang , Zheng Lin , Jun Xu , Wenda Jin , Shao-Ping Lu , Deng-Ping Fan

RGB-D salient object detection (SOD), aiming to highlight prominent regions of a given scene by jointly modeling RGB and depth information, is one of the challenging pixel-level prediction tasks. Recently, the dual-attention mechanism has…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Kang Yi , Haoran Tang , Yumeng Li , Jing Xu , Jun Zhang

Most existing infrared-visible image fusion (IVIF) methods assume high-quality inputs, and therefore struggle to handle dual-source degraded scenarios, typically requiring manual selection and sequential application of multiple…

Computer Vision and Pattern Recognition · Computer Science 2025-09-08 Tianpei Zhang , Jufeng Zhao , Yiming Zhu , Guangmang Cui

We present an effective method to progressively integrate and refine the cross-modality complementarities for RGB-D salient object detection (SOD). The proposed network mainly solves two challenging issues: 1) how to effectively integrate…

Computer Vision and Pattern Recognition · Computer Science 2020-07-15 Chongyi Li , Runmin Cong , Yongri Piao , Qianqian Xu , Chen Change Loy

Semantic segmentation of RGB-D images involves understanding the appearance and spatial relationships of objects within a scene, which requires careful consideration of various factors. However, in indoor environments, the simple input of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-07 Shuai Zhang , Minghong Xie

In this paper, we propose a \textbf{Tr}ansformer-based RGB-D \textbf{e}gocentric \textbf{a}ction \textbf{r}ecognition framework, called Trear. It consists of two modules, inter-frame attention encoder and mutual-attentional fusion block.…

Computer Vision and Pattern Recognition · Computer Science 2021-01-12 Xiangyu Li , Yonghong Hou , Pichao Wang , Zhimin Gao , Mingliang Xu , Wanqing Li

Salient Object Detection is the task of predicting the human attended region in a given scene. Fusing depth information has been proven effective in this task. The main challenge of this problem is how to aggregate the complementary…

Computer Vision and Pattern Recognition · Computer Science 2022-06-08 Chao Zeng , Sam Kwong

Robust object recognition is a crucial ingredient of many, if not all, real-world robotics applications. This paper leverages recent progress on Convolutional Neural Networks (CNNs) and proposes a novel RGB-D architecture for object…

Computer Vision and Pattern Recognition · Computer Science 2015-08-19 Andreas Eitel , Jost Tobias Springenberg , Luciano Spinello , Martin Riedmiller , Wolfram Burgard

Point cloud registration is a task to estimate the rigid transformation between two unaligned scans, which plays an important role in many computer vision applications. Previous learning-based works commonly focus on supervised…

Computer Vision and Pattern Recognition · Computer Science 2023-08-10 Mingzhi Yuan , Kexue Fu , Zhihao Li , Yucong Meng , Manning Wang

With the rapid development of deep learning technology, more and more face forgeries by deepfake are widely spread on social media, causing serious social concern. Face forgery detection has become a research hotspot in recent years, and…

Computer Vision and Pattern Recognition · Computer Science 2021-09-30 Hao Lin , Weiqi Luo , Kangkang Wei , Minglin Liu

The fusion of images taken by heterogeneous sensors helps to enrich the information and improve the quality of imaging. In this article, we present a hybrid model consisting of a convolutional encoder and a Transformer-based decoder to fuse…

Computer Vision and Pattern Recognition · Computer Science 2022-10-19 Yu Yuan , Jiaqi Wu , Zhongliang Jing , Henry Leung , Han Pan

RGB-Thermal (RGB-T) crowd counting is a challenging task, which uses thermal images as complementary information to RGB images to deal with the decreased performance of unimodal RGB-based methods in scenes with low-illumination or similar…

Computer Vision and Pattern Recognition · Computer Science 2022-08-16 Pengyu Chen , Junyu Gao , Yuan Yuan , Qi Wang