English
Related papers

Related papers: RTFDNet: Fusion-Decoupling for Robust RGB-T Segmen…

200 papers

Multimodal information processing has become increasingly important for enhancing image classification performance. However, the intricate and implicit dependencies across different modalities often hinder conventional methods from…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Yang Qiao , Xiaoyu Zhong , Xiaofeng Gu , Zhiguo Yu

Joint detection of drivable areas and road anomalies is very important for mobile robots. Recently, many semantic segmentation approaches based on convolutional neural networks (CNNs) have been proposed for pixel-wise drivable area and road…

Computer Vision and Pattern Recognition · Computer Science 2021-04-21 Hengli Wang , Rui Fan , Yuxiang Sun , Ming Liu

Cross-modal learning has become a fundamental paradigm for integrating heterogeneous information sources such as images, text, and structured attributes. However, multimodal representations often suffer from modality dominance, redundant…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Xuecheng Li , Weikuan Jia , Alisher Kurbonaliev , Qurbonaliev Alisher , Khudzhamkulov Rustam , Ismoilov Shuhratjon , Eshmatov Javhariddin , Yuanjie Zheng

The Vision Transformer (ViT) architecture has established its place in computer vision literature, however, training ViTs for RGB-D object recognition remains an understudied topic, viewed in recent literature only through the lens of…

Computer Vision and Pattern Recognition · Computer Science 2023-03-08 Georgios Tziafas , Hamidreza Kasaei

Focusing on the issue of how to effectively capture and utilize cross-modality information in RGB-D salient object detection (SOD) task, we present a convolutional neural network (CNN) model, named CIR-Net, based on the novel cross-modality…

Computer Vision and Pattern Recognition · Computer Science 2022-11-23 Runmin Cong , Qinwei Lin , Chen Zhang , Chongyi Li , Xiaochun Cao , Qingming Huang , Yao Zhao

Scene understanding plays a critical role in enabling intelligence and autonomy in robotic systems. Traditional approaches often face challenges, including occlusions, ambiguous boundaries, and the inability to adapt attention based on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Guodong Sun , Junjie Liu , Gaoyang Zhang , Bo Wu , Yang Zhang

Existing RGB-D salient object detection (SOD) models usually treat RGB and depth as independent information and design separate networks for feature extraction from each. Such schemes can easily be constrained by a limited amount of…

Computer Vision and Pattern Recognition · Computer Science 2021-04-19 Keren Fu , Deng-Ping Fan , Ge-Peng Ji , Qijun Zhao , Jianbing Shen , Ce Zhu

4D automotive radar is indispensable for autonomous driving due to its low cost and robustness, yet its point cloud sparsity challenges 3D object detection. Existing 4D radar-camera fusion methods focus on complex fusion strategies, trading…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Weiyi Xiong , Bing Zhu

While deep learning, particularly convolutional neural networks (CNNs), has revolutionized remote sensing (RS) change detection (CD), existing approaches often miss crucial features due to neglecting global context and incomplete change…

Multimedia · Computer Science 2024-07-04 Yuhao Gao , Gensheng Pei , Mengmeng Sheng , Zeren Sun , Tao Chen , Yazhou Yao

Guided depth map super-resolution (GDSR), as a hot topic in multi-modal image processing, aims to upsample low-resolution (LR) depth maps with additional information involved in high-resolution (HR) RGB images from the same scene. The…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Zixiang Zhao , Jiangshe Zhang , Xiang Gu , Chengli Tan , Shuang Xu , Yulun Zhang , Radu Timofte , Luc Van Gool

This paper proposes a novel joint learning and densely-cooperative fusion (JL-DCF) architecture for RGB-D salient object detection. Existing models usually treat RGB and depth as independent information and design separate networks for…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Keren Fu , Deng-Ping Fan , Ge-Peng Ji , Qijun Zhao

For both visible and infrared images have their own advantages and disadvantages, RGBT tracking has attracted more and more attention. The key points of RGBT tracking lie in feature extraction and feature fusion of visible and infrared…

Computer Vision and Pattern Recognition · Computer Science 2023-01-12 Jingchao Peng , Haitao Zhao , Zhengwei Hu

Surgical scene understanding is a key technical component for enabling intelligent and context aware systems that can transform various aspects of surgical interventions. In this work, we focus on the semantic segmentation task, propose a…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Muhammad Abdullah Jamal , Omid Mohareri

In this paper we present a new approach for feature fusion between RGB and LWIR Thermal images for the task of semantic segmentation for driving perception. We propose DooDLeNet, a double DeepLab architecture with specialized…

Machine Learning · Computer Science 2022-04-22 Oriel Frigo , Lucien Martin-Gaffé , Catherine Wacongne

Indoor semantic segmentation is fundamental to computer vision and robotics, supporting applications such as autonomous navigation, augmented reality, and smart environments. Although RGB-D fusion leverages complementary appearance and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Yan Gong , Jianli Lu , Yongsheng Gao , Jie Zhao , Xiaojuan Zhang , Susanto Rahardja

In recent years, the research community has shown a lot of interest to panoramic images that offer a 360-degree directional perspective. Multiple data modalities can be fed, and complimentary characteristics can be utilized for more robust…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Suresh Guttikonda , Jason Rambach

Diffusion transformers have demonstrated remarkable generation quality, albeit requiring longer training iterations and numerous inference steps. In each denoising step, diffusion transformers encode the noisy inputs to extract the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Shuai Wang , Zhi Tian , Weilin Huang , Limin Wang

4D millimeter-wave (mmWave) radar has been widely adopted in autonomous driving and robot perception due to its low cost and all-weather robustness. However, point-cloud-based radar representations suffer from information loss due to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Runwei Guan , Jianan Liu , Shaofeng Liang , Fangqiang Ding , Shanliang Yao , Xiaokai Bai , Daizong Liu , Tao Huang , Guoqiang Mao , Hui Xiong

Most existing lightweight RGB-D salient object detection (SOD) models are based on two-stream structure or single-stream structure. The former one first uses two sub-networks to extract unimodal features from RGB and depth images,…

Computer Vision and Pattern Recognition · Computer Science 2022-11-23 Nianchang Huang , Qiang Zhang , Jungong Han

Depth information matters in RGB-D semantic segmentation task for providing additional geometric information to color images. Most existing methods exploit a multi-stage fusion strategy to propagate depth feature to the RGB branch. However,…

Computer Vision and Pattern Recognition · Computer Science 2021-01-27 Sihan Chen , Xinxin Zhu , Wei Liu , Xingjian He , Jing Liu