English
Related papers

Related papers: Multi-Modal Hybrid Learning and Sequential Trainin…

200 papers

Multispectral object detection aims to leverage complementary information from visible (RGB) and infrared (IR) modalities to enable robust performance under diverse environmental conditions. Our key insight, derived from wavelet analysis…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Seongmin Hwang , Daeyoung Han , Moongu Jeon

Focusing on the issue of how to effectively capture and utilize cross-modality information in RGB-D salient object detection (SOD) task, we present a convolutional neural network (CNN) model, named CIR-Net, based on the novel cross-modality…

Computer Vision and Pattern Recognition · Computer Science 2022-11-23 Runmin Cong , Qinwei Lin , Chen Zhang , Chongyi Li , Xiaochun Cao , Qingming Huang , Yao Zhao

Saliency Prediction aims to predict the attention distribution of human eyes given an RGB image. Most of the recent state-of-the-art methods are based on deep image feature representations from traditional CNNs. However, the traditional…

Computer Vision and Pattern Recognition · Computer Science 2023-01-27 Shuo Zhang

Integrating multispectral data in object detection, especially visible and infrared images, has received great attention in recent years. Since visible (RGB) and infrared (IR) images can provide complementary information to handle light…

Computer Vision and Pattern Recognition · Computer Science 2022-09-29 Maoxun Yuan , Yinyan Wang , Xingxing Wei

Visual object tracking, which is primarily based on visible light image sequences, encounters numerous challenges in complicated scenarios, such as low light conditions, high dynamic ranges, and background clutter. To address these…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Hongze Sun , Rui Liu , Wuque Cai , Jun Wang , Yue Wang , Huajin Tang , Yan Cui , Dezhong Yao , Daqing Guo

The RGB-infrared cross-modality person re-identification (ReID) task aims to recognize the images of the same identity between the visible modality and the infrared modality. Existing methods mainly use a two-stream architecture to…

Computer Vision and Pattern Recognition · Computer Science 2021-10-22 Yajun Gao , Tengfei Liang , Yi Jin , Xiaoyan Gu , Wu Liu , Yidong Li , Congyan Lang

The multi-modal salient object detection model based on RGB-D information has better robustness in the real world. However, it remains nontrivial to better adaptively balance effective multi-modal information in the feature fusion phase. In…

Computer Vision and Pattern Recognition · Computer Science 2022-02-09 Jinchao Zhu , Xiaoyu Zhang , Xian Fang , Feng Dong , Qiu Yu

RGB-thermal salient object detection (SOD) aims to segment the common prominent regions of visible image and corresponding thermal infrared image that we call it RGBT SOD. Existing methods don't fully explore and exploit the potentials of…

Computer Vision and Pattern Recognition · Computer Science 2021-07-07 Zhengzheng Tu , Zhun Li , Chenglong Li , Yang Lang , Jin Tang

Gesture recognition is getting more and more popular due to various application possibilities in human-machine interaction. Existing multi-modal gesture recognition systems take multi-modal data as input to improve accuracy, but such…

Computer Vision and Pattern Recognition · Computer Science 2021-11-01 Dinghao Fan , Hengjie Lu , Shugong Xu , Shan Cao

Multimodal fusion frameworks for Human Action Recognition (HAR) using depth and inertial sensor data have been proposed over the years. In most of the existing works, fusion is performed at a single level (feature level or decision level),…

Machine Learning · Computer Science 2019-10-28 Zeeshan Ahmad , Naimul Khan

Self-supervised cross-modal super-resolution (SR) can overcome the difficulty of acquiring paired training data, but is challenging because only low-resolution (LR) source and high-resolution (HR) guide images from different modalities are…

Computer Vision and Pattern Recognition · Computer Science 2022-07-20 Xiaoyu Dong , Naoto Yokoya , Longguang Wang , Tatsumi Uezato

RGB and thermal source data suffer from both shared and specific challenges, and how to explore and exploit them plays a critical role to represent the target appearance in RGBT tracking. In this paper, we propose a novel challenge-aware…

Computer Vision and Pattern Recognition · Computer Science 2020-07-28 Chenglong Li , Lei Liu , Andong Lu , Qing Ji , Jin Tang

Heterogeneous data modalities can provide complementary cues for several tasks, usually leading to more robust algorithms and better performance. However, while training data can be accurately collected to include a variety of sensory…

Computer Vision and Pattern Recognition · Computer Science 2019-07-29 Nuno C. Garcia , Pietro Morerio , Vittorio Murino

Co-saliency detection aims at extracting the common salient regions from an image group containing two or more relevant images. It is a newly emerging topic in computer vision community. Different from the most existing co-saliency methods…

Computer Vision and Pattern Recognition · Computer Science 2017-11-22 Runmin Cong , Jianjun Lei , Huazhu Fu , Qingming Huang , Xiaochun Cao , Chunping Hou

Light field data exhibit favorable characteristics conducive to saliency detection. The success of learning-based light field saliency detection is heavily dependent on how a comprehensive dataset can be constructed for higher…

Computer Vision and Pattern Recognition · Computer Science 2021-01-01 Yongri Piao , Zhengkun Rong , Shuang Xu , Miao Zhang , Huchuan Lu

Multimodal (e.g., RGB-Depth/RGB-Thermal) fusion has shown great potential for improving semantic segmentation in complex scenes (e.g., indoor/low-light conditions). Existing approaches often fully fine-tune a dual-branch encoder-decoder…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Shaohua Dong , Yunhe Feng , Qing Yang , Yan Huang , Dongfang Liu , Heng Fan

Multimodal deep learning, especially vision-language models, have gained significant traction in recent years, greatly improving performance on many downstream tasks, including content moderation and violence detection. However, standard…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Zhuokai Zhao , Harish Palani , Tianyi Liu , Lena Evans , Ruth Toner

This paper introduces a novel deep learning-based multimodal fusion architecture aimed at enhancing the perception capabilities of autonomous navigation robots in complex environments. By utilizing innovative feature extraction modules,…

Machine Learning · Computer Science 2025-04-29 Delun Lai , Yeyubei Zhang , Yunchong Liu , Chaojie Li , Huadong Mo

Scene understanding based on image segmentation is a crucial component of autonomous vehicles. Pixel-wise semantic segmentation of RGB images can be advanced by exploiting complementary features from the supplementary modality (X-modality).…

Computer Vision and Pattern Recognition · Computer Science 2023-11-27 Jiaming Zhang , Huayao Liu , Kailun Yang , Xinxin Hu , Ruiping Liu , Rainer Stiefelhagen

Multi-Modal Object Detection (MMOD), due to its stronger adaptability to various complex environments, has been widely applied in various applications. Extensive research is dedicated to the RGB-IR object detection, primarily focusing on…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Tianyi Zhao , Boyang Liu , Yanglei Gao , Yiming Sun , Maoxun Yuan , Xingxing Wei