中文
相关论文

相关论文: Bridging the RGB-IR Gap: Consensus and Discrepancy…

200 篇论文

RGB-thermal salient object detection (RGB-T SOD) aims to locate the common prominent objects of an aligned visible and thermal infrared image pair and accurately segment all the pixels belonging to those objects. It is promising in…

计算机视觉与模式识别 · 计算机科学 2022-07-11 Xiurong Jiang , Lin Zhu , Yifan Hou , Hui Tian

Enhancing scene understanding in adverse visibility conditions remains a critical challenge for surveillance and autonomous navigation systems. Conventional imaging modalities, such as RGB and thermal infrared (MWIR / LWIR), when fused,…

机器学习 · 计算机科学 2025-11-25 Muhammad Ishfaq Hussain , Ma Van Linh , Zubia Naz , Unse Fatima , Yeongmin Ko , Moongu Jeon

RGB-D saliency detection aims to fuse multi-modal cues to accurately localize salient regions. Existing works often adopt attention modules for feature modeling, with few methods explicitly leveraging fine-grained details to merge with…

计算机视觉与模式识别 · 计算机科学 2023-04-19 Zongwei Wu , Guillaume Allibert , Fabrice Meriaudeau , Chao Ma , Cédric Demonceaux

Transparent object perception remains a major challenge in computer vision research, as transparency confounds both depth estimation and semantic segmentation. Recent work has explored multi-task learning frameworks to improve robustness,…

计算机视觉与模式识别 · 计算机科学 2025-11-20 Gbenga Omotara , Ramy Farag , Seyed Mohamad Ali Tousi , G. N. DeSouza

Image-text matching (ITM) is a fundamental problem in computer vision. The key issue lies in jointly learning the visual and textual representation to estimate their similarity accurately. Most existing methods focus on feature enhancement…

计算机视觉与模式识别 · 计算机科学 2024-06-28 Xuri Ge , Fuhai Chen , Songpei Xu , Fuxiang Tao , Jie Wang , Joemon M. Jose

Efficiently exploiting multi-modal inputs for accurate RGB-D saliency detection is a topic of high interest. Most existing works leverage cross-modal interactions to fuse the two streams of RGB-D for intermediate features' enhancement. In…

计算机视觉与模式识别 · 计算机科学 2022-08-31 Zongwei Wu , Shriarulmozhivarman Gobichettipalayam , Brahim Tamadazte , Guillaume Allibert , Danda Pani Paudel , Cédric Demonceaux

Data-fusion networks have shown significant promise for RGB-thermal scene parsing. However, the majority of existing studies have relied on symmetric duplex encoders for heterogeneous feature extraction and fusion, paying inadequate…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Jiahang Li , Peng Yun , Yang Xu , Ye Zhang , Mingjian Sun , Qijun Chen , Ilin Alexander , Rui Fan

RGB-Thermal (T) crowd counting aims to integrate visible-spectrum and thermal infrared information to improve the robustness of crowd density estimation in complex scenes. Although existing studies generally improve counting accuracy…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Jinghao Shi , Mengqi Lei , Kunliang He , Yun Li , Wei Bao , Siqi Li

Recent advancements in 3D scene understanding have made significant strides in enabling interaction with scenes using open-vocabulary queries, particularly for VR/AR and robotic applications. Nevertheless, existing methods are hindered by…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Dianyi Yang , Xihan Wang , Yu Gao , Shiyang Liu , Bohan Ren , Yufeng Yue , Yi Yang

Event cameras capture microsecond-level motion cues that complement RGB sensors. However, the prevailing paradigm of treating RGB-Event perception as a fusion problem is ill-posed, as it ignores the intrinsic (i) Spatiotemporal and (ii)…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Zhen Yao , Xiaowen Ying , Zhiyu Zhu , Mooi Choo Chuah

Technological development aims to produce generations of increasingly efficient robots able to perform complex tasks. This requires considerable efforts, from the scientific community, to find new algorithms that solve computer vision…

计算机视觉与模式识别 · 计算机科学 2018-09-06 Mirco Planamente , Mohammad Reza Loghmani , Barbara Caputo

Multispectral image fusion is a computer vision process that is essential to remote sensing. For applications such as dehazing and object detection, there is a need to offer solutions that can perform in real-time on any type of scene.…

计算机视觉与模式识别 · 计算机科学 2023-05-09 Nati Ofir , Jean-Christophe Nebel

Feature modeling of different modalities is a basic problem in current research of cross-modal information retrieval. Existing models typically project texts and images into one embedding space, in which semantically similar information…

多媒体 · 计算机科学 2019-06-13 Jing Yu , Chenghao Yang , Zengchang Qin , Zhuoqian Yang , Yue Hu , Weifeng Zhang

2D RGB images and 3D LIDAR point clouds provide complementary knowledge for the perception system of autonomous vehicles. Several 2D and 3D fusion methods have been explored for the LIDAR semantic segmentation task, but they suffer from…

计算机视觉与模式识别 · 计算机科学 2023-07-11 Jun Cen , Shiwei Zhang , Yixuan Pei , Kun Li , Hang Zheng , Maochun Luo , Yingya Zhang , Qifeng Chen

Glass surfaces are becoming increasingly ubiquitous as modern buildings tend to use a lot of glass panels. This, however, poses substantial challenges to the operations of autonomous systems such as robots, self-driving cars, and drones, as…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Jiaying Lin , Yuen-Hei Yeung , Shuquan Ye , Rynson W. H. Lau

Efficient RGB-D semantic segmentation has received considerable attention in mobile robots, which plays a vital role in analyzing and recognizing environmental information. According to previous studies, depth information can provide…

计算机视觉与模式识别 · 计算机科学 2023-08-14 Yang Zhang , Chenyun Xiong , Junjie Liu , Xuhui Ye , Guodong Sun

Geometric information in the normalized digital surface models (nDSM) is highly correlated with the semantic class of the land cover. Exploiting two modalities (RGB and nDSM (height)) jointly has great potential to improve the segmentation…

计算机视觉与模式识别 · 计算机科学 2023-05-25 Zhitong Xiong , Sining Chen , Yi Wang , Lichao Mou , Xiao Xiang Zhu

Image-text matching is a key multimodal task that aims to model the semantic association between images and text as a matching relationship. With the advent of the multimedia information age, image, and text data show explosive growth, and…

机器学习 · 计算机科学 2024-06-24 Jinyin Wang , Haijing Zhang , Yihao Zhong , Yingbin Liang , Rongwei Ji , Yiru Cang

The performance of traditional text-image person retrieval task is easily affected by lighting variations due to imaging limitations of visible spectrum sensors. In recent years, cross-modal information fusion has emerged as an effective…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Yifei Deng , Chenglong Li , Zhenyu Chen , Zihen Xu , Jin Tang

Existing RGB-D saliency detection models do not explicitly encourage RGB and depth to achieve effective multi-modal learning. In this paper, we introduce a novel multi-stage cascaded learning framework via mutual information minimization to…

计算机视觉与模式识别 · 计算机科学 2022-01-07 Jing Zhang , Deng-Ping Fan , Yuchao Dai , Xin Yu , Yiran Zhong , Nick Barnes , Ling Shao