English
Related papers

Related papers: Complementary Random Masking for RGB-Thermal Seman…

200 papers

Multimodal Dataset Distillation (MDD) seeks to condense large-scale image-text datasets into compact surrogates while retaining their effectiveness for cross-modal learning. Despite recent progress, existing MDD approaches often suffer from…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Xin Zhang , Ziruo Zhang , Jiawei Du , Zuozhu Liu , Joey Tianyi Zhou

In the RGB-D vision community, extensive research has been focused on designing multi-modal learning strategies and fusion structures. However, the complementary and fusion mechanisms in RGB-D models remain a black box. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Hao Chen , Haoran Zhou , Yunshu Zhang , Zheng Lin , Yongjian Deng

Future advancements in robot autonomy and sophistication of robotics tasks rest on robust, efficient, and task-dependent semantic understanding of the environment. Semantic segmentation is the problem of simultaneous segmentation and…

Computer Vision and Pattern Recognition · Computer Science 2016-06-06 Md. Alimoor Reza , Jana Kosecka

This paper introduces a novel unified representation of diffusion models for image generation and segmentation. Specifically, we use a colormap to represent entity-level masks, addressing the challenge of varying entity numbers while…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Lu Qi , Lehan Yang , Weidong Guo , Yu Xu , Bo Du , Varun Jampani , Ming-Hsuan Yang

Multispectral object detection, utilizing both visible (RGB) and thermal infrared (T) modals, has garnered significant attention for its robust performance across diverse weather and lighting conditions. However, effectively exploiting the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Jinzhong Wang , Xuetao Tian , Shun Dai , Tao Zhuo , Haorui Zeng , Hongjuan Liu , Jiaqi Liu , Xiuwei Zhang , Yanning Zhang

The integration of RGB and depth modalities significantly enhances the accuracy of segmenting complex indoor scenes, with depth data from RGB-D cameras playing a crucial role in this improvement. However, collecting an RGB-D dataset is more…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Xinhua Xu , Hong Liu , Jianbing Wu , Jinfu Liu

RGB-thermal (RGB-T) semantic segmentation improves the environmental perception of autonomous platforms in challenging conditions. Prevailing models employ encoders pre-trained on RGB images to extract features from both RGB and infrared…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Xiaodong Guo , Tong Liu , Yike Li , Zi'ang Lin , Zhihong Deng

Developing robust multi-modal feature representations is crucial for enhancing object tracking performance. In pursuit of this objective, a novel X Modality Assisting Network (X-Net) is introduced, which explores the impact of the fusion…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Zhaisheng Ding , Haiyan Li , Ruichao Hou , Yanyu Liu , Shidong Xie

In this paper, we tackle the problem of RGB-D semantic segmentation of indoor images. We take advantage of deconvolutional networks which can predict pixel-wise class labels, and develop a new structure for deconvolution of multiple…

Computer Vision and Pattern Recognition · Computer Science 2016-08-04 Jinghua Wang , Zhenhua Wang , Dacheng Tao , Simon See , Gang Wang

How to effectively fuse cross-modal information is the key problem for RGB-D salient object detection. Early fusion and the result fusion schemes fuse RGB and depth information at the input and output stages, respectively, hence incur the…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Nian Liu , Ni Zhang , Ling Shao , Junwei Han

RGB-T semantic segmentation is a key technique for autonomous driving scenes understanding. For the existing RGB-T semantic segmentation methods, however, the effective exploration of the complementary relationship between different…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Ying Lv , Zhi Liu , Gongyang Li

RGB-Thermal (RGBT) multispectral vision is essential for robust perception in complex environments. Most RGBT tasks follow a case-by-case research paradigm, relying on manually customized models to learn task-oriented representations.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Kailai Zhou , Fuqiang Yang , Shixian Wang , Bihan Wen , Chongde Zi , Linsen Chen , Qiu Shen , Xun Cao

Paired RGB-thermal data is crucial for visual-thermal sensor fusion and cross-modality tasks, including important applications such as multi-modal image alignment and retrieval. However, the scarcity of synchronized and calibrated…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Jiuhong Xiao , Roshan Nayak , Ning Zhang , Daniel Tortei , Giuseppe Loianno

Multimodal remote sensing data provide complementary information for semantic segmentation, but in real-world deployments, some modalities may be unavailable due to sensor failures, acquisition issues, or challenging atmospheric conditions.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Irem Ulku , Erdem Akagündüz , Ömer Özgür Tanrıöver

Accurate segmentation of brain images typically requires the integration of complementary information from multiple image modalities. However, clinical data for all modalities may not be available for every patient, creating a significant…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Haitao Li , Ziyu Li , Yiheng Mao , Zhengyao Ding , Zhengxing Huang

Semantic segmentation has made striking progress due to the success of deep convolutional neural networks. Considering the demands of autonomous driving, real-time semantic segmentation has become a research hotspot these years. However,…

Computer Vision and Pattern Recognition · Computer Science 2020-06-30 Lei Sun , Kailun Yang , Xinxin Hu , Weijian Hu , Kaiwei Wang

Multi-modal image segmentation faces real-world deployment challenges from incomplete/corrupted modalities degrading performance. While existing methods address training-inference modality gaps via specialized per-combination models, they…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Xiaoqi Zhao , Youwei Pang , Chenyang Yu , Lihe Zhang , Huchuan Lu , Shijian Lu , Georges El Fakhri , Xiaofeng Liu

Medical image segmentation of tumors and organs at risk is a time-consuming yet critical process in the clinic that utilizes multi-modality imaging (e.g, different acquisitions, data types, and sequences) to increase segmentation precision.…

Image and Video Processing · Electrical Eng. & Systems 2023-06-07 Qisheng He , Nicholas Summerfield , Ming Dong , Carri Glide-Hurst

In this work, we address the problem how a network for action recognition that has been trained on a modality like RGB videos can be adapted to recognize actions for another modality like sequences of 3D human poses. To this end, we extract…

Computer Vision and Pattern Recognition · Computer Science 2019-10-11 Fida Mohammad Thoker , Juergen Gall

Efficient and robust 3D scene representation is crucial in autonomous driving, robotics, and related fields. While RGB images provide valuable content for 3D reconstruction, other modalities like thermal or depth can enable additional…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Manoj Biswanath , Chenxin Cai , Hannah Schieber , Daniel Roth , Benjamin Busam
‹ Prev 1 3 4 5 6 7 10 Next ›