English
Related papers

Related papers: Transformer-based Network for RGB-D Saliency Detec…

200 papers

Task-specific data-fusion networks have marked considerable achievements in urban scene parsing. Among these networks, our recently proposed RoadFormer successfully extracts heterogeneous features from RGB images and surface normal maps and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-23 Jianxin Huang , Jiahang Li , Ning Jia , Yuxiang Sun , Chengju Liu , Qijun Chen , Rui Fan

Existing Transformer-based RGBT tracking methods either use cross-attention to fuse the two modalities, or use self-attention and cross-attention to model both modality-specific and modality-sharing information. However, the significant…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Yabin Zhu , Chenglong Li , Xiao Wang , Jin Tang , Zhixiang Huang

Co-saliency detection aims at extracting the common salient regions from an image group containing two or more relevant images. It is a newly emerging topic in computer vision community. Different from the most existing co-saliency methods…

Computer Vision and Pattern Recognition · Computer Science 2017-11-22 Runmin Cong , Jianjun Lei , Huazhu Fu , Qingming Huang , Xiaochun Cao , Chunping Hou

Depth information matters in RGB-D semantic segmentation task for providing additional geometric information to color images. Most existing methods exploit a multi-stage fusion strategy to propagate depth feature to the RGB branch. However,…

Computer Vision and Pattern Recognition · Computer Science 2021-01-27 Sihan Chen , Xinxin Zhu , Wei Liu , Xingjian He , Jing Liu

Bottom-up and top-down visual cues are two types of information that helps the visual saliency models. These salient cues can be from spatial distributions of the features (space-based saliency) or contextual / task-dependent features…

Computer Vision and Pattern Recognition · Computer Science 2018-07-05 Nevrez Imamoglu , Wataru Shimoda , Chi Zhang , Yuming Fang , Asako Kanezaki , Keiji Yanai , Yoshifumi Nishida

The main problem in RGB-T tracking is the correct and optimal merging of the cross-modal features of visible and thermal images. Some previous methods either do not fully exploit the potential of RGB and TIR information for channel and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Yunfeng Li , Bo Wang , Ye Li

Mapping and localization are essential capabilities of robotic systems. Although the majority of mapping systems focus on static environments, the deployment in real-world situations requires them to handle dynamic objects. In this paper,…

Robotics · Computer Science 2019-08-30 Emanuele Palazzolo , Jens Behley , Philipp Lottes , Philippe Giguère , Cyrill Stachniss

Multispectral image pairs can provide the combined information, making object detection applications more reliable and robust in the open world. To fully exploit the different modalities, we present a simple yet effective cross-modality…

Image and Video Processing · Electrical Eng. & Systems 2022-10-05 Fang Qingyun , Han Dapeng , Wang Zhaokui

Event camera-based pattern recognition is a newly arising research topic in recent years. Current researchers usually transform the event streams into images, graphs, or voxels, and adopt deep neural networks for event-based classification.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Xiao Wang , Yao Rong , Zongzhen Wu , Lin Zhu , Bo Jiang , Jin Tang , Yonghong Tian

We address the problem of glass surface segmentation with an RGB-D camera, with a focus on effectively fusing RGB and depth information. To this end, we propose a Weighted Feature Fusion (WFF) module that dynamically and adaptively combines…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Henghong Lin , Zihan Zhu , Tao Wang , Anastasia Ioannou , Yuanshui Huang

Visual saliency detection model simulates the human visual system to perceive the scene, and has been widely used in many vision tasks. With the acquisition technology development, more comprehensive information, such as depth cue,…

Computer Vision and Pattern Recognition · Computer Science 2019-09-04 Runmin Cong , Jianjun Lei , Huazhu Fu , Ming-Ming Cheng , Weisi Lin , Qingming Huang

In this paper, we tackle the problem of RGB-D semantic segmentation of indoor images. We take advantage of deconvolutional networks which can predict pixel-wise class labels, and develop a new structure for deconvolution of multiple…

Computer Vision and Pattern Recognition · Computer Science 2016-08-04 Jinghua Wang , Zhenhua Wang , Dacheng Tao , Simon See , Gang Wang

In RGB-D semantic segmentation for indoor scenes, a key challenge is effectively integrating the rich color information from RGB images with the spatial distance information from depth images. However, most existing methods overlook the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Shuobin Wei , Zhuang Zhou , Zhengan Lu , Zizhao Yuan , Binghua Su

Survival prediction plays a crucial role in assisting clinicians with the development of cancer treatment protocols. Recent evidence shows that multimodal data can help in the diagnosis of cancer disease and improve survival prediction.…

Image and Video Processing · Electrical Eng. & Systems 2023-11-14 Ruiquan Ge , Xiangyang Hu , Rungen Huang , Gangyong Jia , Yaqi Wang , Renshu Gu , Changmiao Wang , Elazab Ahmed , Linyan Wang , Juan Ye , Ye Li

Deep convolutional neural network (CNN) based salient object detection methods have achieved state-of-the-art performance and outperform those unsupervised methods with a wide margin. In this paper, we propose to integrate deep and…

Computer Vision and Pattern Recognition · Computer Science 2017-06-05 Jing Zhang , Bo Li , Yuchao Dai , Fatih Porikli , Mingyi He

While deep learning, particularly convolutional neural networks (CNNs), has revolutionized remote sensing (RS) change detection (CD), existing approaches often miss crucial features due to neglecting global context and incomplete change…

Multimedia · Computer Science 2024-07-04 Yuhao Gao , Gensheng Pei , Mengmeng Sheng , Zeren Sun , Tao Chen , Yazhou Yao

Recently, CNN and Transformer hybrid networks demonstrated excellent performance in face super-resolution (FSR) tasks. Since numerous features at different scales in hybrid networks, how to fuse these multiscale features and promote their…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Xujie Wan , Wenjie Li , Guangwei Gao , Huimin Lu , Jian Yang , Chia-Wen Lin

Providing machines with the ability to recognize objects like humans has always been one of the primary goals of machine vision. The introduction of RGB-D cameras has paved the way for a significant leap forward in this direction thanks to…

Computer Vision and Pattern Recognition · Computer Science 2019-02-26 Mohammad Reza Loghmani , Mirco Planamente , Barbara Caputo , Markus Vincze

Traditional systems typically require different models for processing different modalities, such as one model for RGB images and another for depth images. Recent research has demonstrated that a single model for one modality can be adapted…

Computer Vision and Pattern Recognition · Computer Science 2023-05-09 Xiaoke Shen , Ioannis Stamos

Visible-infrared object detection has gained sufficient attention due to its detection performance in low light, fog, and rain conditions. However, visible and infrared modalities captured by different sensors exist the information…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Wencong Wu , Xiuwei Zhang , Hanlin Yin , Shun Dai , Hongxi Zhang , Yanning Zhang