中文
相关论文

相关论文: Missing No More: Dictionary-Guided Cross-Modal Ima…

200 篇论文

Composed image retrieval (CIR) aims to retrieve the target image based on a multimodal query, i.e., a reference image paired with corresponding modification text. Recent CIR studies leverage vision-language pre-trained (VLP) methods as the…

多媒体 · 计算机科学 2024-04-25 Haokun Wen , Xuemeng Song , Xiaolin Chen , Yinwei Wei , Liqiang Nie , Tat-Seng Chua

Recent advances in multimodal recommendation (MMR) highlight the potential of integrating visual and textual content to enrich item representations. However, existing methods often rely on coarse visual features and naive fusion strategies,…

信息检索 · 计算机科学 2025-11-11 Hai-Dang Kieu , Min Xu , Thanh Trung Huynh , Dung D. Le

Image fusion, a fundamental low-level vision task, aims to integrate multiple image sequences into a single output while preserving as much information as possible from the input. However, existing methods face several significant…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Zihan Cao , Yu Zhong , Ziqi Wang , Liang-Jian Deng

Multispectral image fusion is a computer vision process that is essential to remote sensing. For applications such as dehazing and object detection, there is a need to offer solutions that can perform in real-time on any type of scene.…

计算机视觉与模式识别 · 计算机科学 2023-05-09 Nati Ofir , Jean-Christophe Nebel

Composed Video Retrieval (CoVR) facilitates video retrieval by combining visual and textual queries. However, existing CoVR frameworks typically fuse multimodal inputs in a single stage, achieving only marginal gains over initial baseline.…

计算机视觉与模式识别 · 计算机科学 2026-01-26 Yuqian Zheng , Mariana-Iuliana Georgescu

The core of video-based visible-infrared person re-identification (VVI-ReID) lies in learning sequence-level modal-invariant representations across different modalities. Recent research tends to use modality-shared language prompts…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Xiaomei Yang , Xizhan Gao , Antai Liu , Kang Wei , Fa Zhu , Guang Feng , Xiaofeng Qu , Sijie Niu

Composed image retrieval (CIR) aims to retrieve a target image that depicts a reference image modified by a textual description. While recent vision-language models (VLMs) achieve promising CIR performance by embedding images and text into…

计算机视觉与模式识别 · 计算机科学 2026-04-13 François Gardères , Camille-Sovanneary Gauthier , Jean Ponce , Shizhe Chen

In this paper, we present a novel deep learning architecture for infrared and visible images fusion problem. In contrast to conventional convolutional networks, our encoding network is combined by convolutional layers, fusion layer and…

计算机视觉与模式识别 · 计算机科学 2019-01-23 Hui Li , Xiao-Jun Wu

The fusion of infrared and visible images is essential in remote sensing applications, as it combines the thermal information of infrared images with the detailed texture of visible images for more accurate analysis in tasks like…

图像与视频处理 · 电气工程与系统科学 2025-06-19 Han Yan , Songlei Xiong , Long Wang , Lihua Jian , Gemine Vivone

Infrared-visible image fusion methods aim at generating fused images with good visual quality and also facilitate the performance of high-level tasks. Indeed, existing semantic-driven methods have considered semantic information injection…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Liying Wang , Xiaoli Zhang , Chuanmin Jia , Siwei Ma

Composed Image Retrieval (CIR) allows users to search for images by combining a reference image with a text prompt that describes desired modifications. While vision-language models like CLIP have popularized this task by embedding multiple…

Infrared and visible image fusion (IVIF) is essential for integrating thermal saliency with textural details to support downstream perception. However, most existing approaches suffer from "semantic blindness," leading to the erroneous…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Xiaoyang Zhang , jinjiang Li , Guodong Fan , Yakun Ju , Linwei Fan , Jun Liu , Alex C. Kot

Existing infrared and visible image fusion methods often face the dilemma of balancing modal information. Generative fusion methods reconstruct fused images by learning from data distributions, but their generative capabilities remain…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Lin Guo , Xiaoqing Luo , Wei Xie , Zhancheng Zhang , Hui Li , Rui Wang , Zhenhua Feng , Xiaoning Song

Infrared and visible image fusion has emerged as a prominent research area in computer vision. However, little attention has been paid to the fusion task in complex scenes, leading to sub-optimal results under interference. To fill this…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Xilai Li , Xiaosong Li , Tianshu Tan , Huafeng Li , Tao Ye

Multi-focus image fusion aims to combine multiple partially focused images into a single all-in-focus image. Although deep learning has shown promise in this task, its effectiveness is often limited by the scarcity of suitable training…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Huangxing Lin , Rongrong Ma , Cheng Wang

Multimodal visual information fusion aims to integrate the multi-sensor data into a single image which contains more complementary information and less redundant features. However the complementary information is hard to extract, especially…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Hui Li , Xiao-Jun Wu

Many image restoration (IR) tasks require both pixel-level fidelity and high-level semantic understanding to recover realistic photos with fine-grained details. However, previous approaches often struggle to effectively leverage both the…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Cuixin Yang , Rongkang Dong , Kin-Man Lam

Recent multimodal retrieval methods have endowed text-based retrievers with multimodal capabilities by utilizing pre-training strategies for visual-text alignment. They often directly fuse the two modalities for cross-reference during the…

计算机视觉与模式识别 · 计算机科学 2025-05-22 Yeong-Joon Ju , Ho-Joong Kim , Seong-Whan Lee

Image fusion aims to combine information from different source images to create a comprehensively representative image. Existing fusion methods are typically helpless in dealing with degradations in low-quality source images and…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Xunpeng Yi , Han Xu , Hao Zhang , Linfeng Tang , Jiayi Ma

Existing text-driven infrared and visible image fusion approaches often rely on textual information at the sentence level, which can lead to semantic noise from redundant text and fail to fully exploit the deeper semantic value of textual…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Wenyu Shao , Hongbo Liu , Yunchuan Ma , Ruili Wang