English
Related papers

Related papers: Multi-Modal Hybrid Learning and Sequential Trainin…

200 papers

Multimodal Magnetic Resonance Imaging (MRI) provides essential complementary information for analyzing brain tumor subregions. While methods using four common MRI modalities for automatic segmentation have shown success, they often face…

Image and Video Processing · Electrical Eng. & Systems 2024-11-14 Runze Cheng , Zhongao Sun , Ye Zhang , Chun Li

RGB-D SOD uses depth information to handle challenging scenes and obtain high-quality saliency maps. Existing state-of-the-art RGB-D saliency detection methods overwhelmingly rely on the strategy of directly fusing depth information.…

Computer Vision and Pattern Recognition · Computer Science 2022-07-12 Xingzhao Jia , Dongye Changlei , Yanjun Peng

Albeit intensively studied, false prediction and unclear boundaries are still major issues of salient object detection. In this paper, we propose a Region Refinement Network (RRN), which recurrently filters redundant information and…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Zhuotao Tian , Hengshuang Zhao , Michelle Shu , Jiaze Wang , Ruiyu Li , Xiaoyong Shen , Jiaya Jia

Salient object detection(SOD) aims at locating the most significant object within a given image. In recent years, great progress has been made in applying SOD on many vision tasks. The depth map could provide additional spatial prior and…

Computer Vision and Pattern Recognition · Computer Science 2021-06-09 Guangyu Ren , Yanchu Xie , Tianhong Dai , Tania Stathaki

Detecting hidden or partially concealed objects remains a fundamental challenge in multimodal environments, where factors like occlusion, camouflage, and lighting variations significantly hinder performance. Traditional RGB-based detection…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Harris Song , Tuan-Anh Vu , Sanjith Menon , Sriram Narasimhan , M. Khalid Jawed

Multi-modal semantic segmentation (MMSS) faces significant challenges in real-world applications due to incomplete, degraded, or missing sensor data. While current MMSS methods typically use self-distillation with modality dropout to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Jiaqi Tan , Xu Zheng , Yang Liu

Training deep models for RGB-D salient object detection (SOD) often requires a large number of labeled RGB-D images. However, RGB-D data is not easily acquired, which limits the development of RGB-D SOD techniques. To alleviate this issue,…

Image and Video Processing · Electrical Eng. & Systems 2022-01-04 Xiaoqiang Wang , Lei Zhu , Siliang Tang , Huazhu Fu , Ping Li , Fei Wu , Yi Yang , Yueting Zhuang

Large-scale pre-trained Vision-Language Models (VLMs) have significantly advanced transfer learning across diverse tasks. However, adapting these models with limited few-shot data often leads to overfitting, undermining their ability to…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Yuncheng Guo , Xiaodong Gu

Depth images and thermal images contain the spatial geometry information and surface temperature information, which can act as complementary information for the RGB modality. However, the quality of the depth and thermal images is often…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Liuxin Bao , Xiaofei Zhou , Xiankai Lu , Yaoqi Sun , Haibing Yin , Zhenghui Hu , Jiyong Zhang , Chenggang Yan

Multi-modal salient object detection (MSOD) aims to boost saliency detection performance by integrating visible sources with depth or thermal infrared ones. Existing methods generally design different fusion schemes to handle certain issues…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Kunpeng Wang , Zhengzheng Tu , Chenglong Li , Cheng Zhang , Bin Luo

RGB-Thermal (RGB-T) pedestrian detection aims to locate the pedestrians in RGB-T image pairs to exploit the complementation between the two modalities for improving detection robustness in extreme conditions. Most existing algorithms assume…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Chao Tian , Zikun Zhou , Yuqing Huang , Gaojun Li , Zhenyu He

In pervasive machine learning, especially in Human Behavior Analysis (HBA), RGB has been the primary modality due to its accessibility and richness of information. However, linked with its benefits are challenges, including sensitivity to…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Christian Stippel , Thomas Heitzinger , Rafael Sterzinger , Martin Kampel

The majority of learning-based semantic segmentation methods are optimized for daytime scenarios and favorable lighting conditions. Real-world driving scenarios, however, entail adverse environmental conditions such as nighttime…

Computer Vision and Pattern Recognition · Computer Science 2020-03-11 Johan Vertens , Jannik Zürn , Wolfram Burgard

Although multimodal large language models (MLLMs) excel in high-level vision-language reasoning, they lack inherent awareness of visual saliency, making it difficult to identify key visual elements. To bridge this gap, we propose…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Long Li , Shuichen Ji , Ziyang Luo , Zhihui Li , Dingwen Zhang , Junwei Han , Nian Liu

Deep learning-based detection networks have made remarkable progress in autonomous driving systems (ADS). ADS should have reliable performance across a variety of ambient lighting and adverse weather conditions. However, luminance…

Computer Vision and Pattern Recognition · Computer Science 2021-11-10 Shruthi Gowda , Bahram Zonooz , Elahe Arani

RGB-T semantic segmentation has been widely adopted to handle hard scenes with poor lighting conditions by fusing different modality features of RGB and thermal images. Existing methods try to find an optimal fusion feature for…

Computer Vision and Pattern Recognition · Computer Science 2023-07-18 Baihong Lin , Zengrong Lin , Yulan Guo , Yulan Zhang , Jianxiao Zou , Shicai Fan

Existing RGB-Event detection methods process the low-information regions of both modalities (background in images and non-event regions in event data) uniformly during feature extraction and fusion, resulting in high computational costs and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Nan Yang , Yang Wang , Zhanwen Liu , Yuchao Dai , Yang Liu , Xiangmo Zhao

In the last decade, the computer vision field has seen significant progress in multimodal data fusion and learning, where multiple sensors, including depth, infrared, and visual, are used to capture the environment across diverse spectral…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Martin Brenner , Napoleon H. Reyes , Teo Susnjak , Andre L. C. Barczak

Hyperspectral imaging can help better understand the characteristics of different materials, compared with traditional image systems. However, only high-resolution multispectral (HrMS) and low-resolution hyperspectral (LrHS) images can…

Computer Vision and Pattern Recognition · Computer Science 2019-01-11 Qi Xie , Minghao Zhou , Qian Zhao , Deyu Meng , Wangmeng Zuo , Zongben Xu

Robust semantic segmentation of road scenes under adverse illumination, lighting, and shadow conditions remain a core challenge for autonomous driving applications. RGB-Thermal fusion is a standard approach, yet existing methods apply…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Ruturaj Reddy , Hrishav Bakul Barua , Junn Yong Loo , Thanh Thi Nguyen , Ganesh Krishnasamy