English
Related papers

Related papers: Cross-modal Full-mode Fine-grained Alignment for T…

200 papers

Text-based person anomaly retrieval has emerged as a challenging task, with most existing approaches relying on complex deep-learning techniques. This raises a research question: How can the model be optimized to achieve greater…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Tien-Huy Nguyen , Huu-Loc Tran , Huu-Phong Phan-Nguyen , Quang-Vinh Dinh

The challenge of Multimodal Deformable Image Registration (MDIR) lies in the conversion and alignment of features between images of different modalities. Generative models (GMs) cannot retain the necessary information enough from the source…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Mingrui Ma , Weijie Wang , Jie Ning , Jianfeng He , Nicu Sebe , Bruno Lepri

We address the problem of visible-infrared person re-identification (VI-reID), that is, retrieving a set of person images, captured by visible or infrared cameras, in a cross-modal setting. Two main challenges in VI-reID are intra-class…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Hyunjong Park , Sanghoon Lee , Junghyup Lee , Bumsub Ham

Image-text retrieval (ITR) is a challenging task in the field of multimodal information processing due to the semantic gap between different modalities. In recent years, researchers have made great progress in exploring the accurate…

Computer Vision and Pattern Recognition · Computer Science 2022-12-19 Jie Guo , Meiting Wang , Yan Zhou , Bin Song , Yuhao Chi , Wei Fan , Jianglong Chang

Weakly supervised text-to-person image matching, as a crucial approach to reducing models' reliance on large-scale manually labeled samples, holds significant research value. However, existing methods struggle to predict complex one-to-many…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Yafei Zhang , Yongle Shang , Huafeng Li

Recent advancements in adapting vision-language pre-training models like CLIP for person re-identification (ReID) tasks often rely on complex adapter design or modality-specific tuning while neglecting cross-modal interaction, leading to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Yunfei Xie , Yuxuan Cheng , Juncheng Wu , Haoyu Zhang , Yuyin Zhou , Shoudong Han

Text-to-image diffusion models often face a severe trilemma in human portrait generation: text-image alignment, photorealism, and human-perceived aesthetics inherently inhibit one another. Supervised Fine-Tuning (SFT) is an effective method…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Yunlong Wang , Jinjin Shi , Wenbin Gao , Xuran Xu , Runyu Shi , Ying Huang

Text-based person search aims to retrieve the corresponding person images in an image database by virtue of a describing sentence about the person, which poses great potential for various applications such as video surveillance. Extracting…

Computer Vision and Pattern Recognition · Computer Science 2019-11-28 Ya Jing , Chenyang Si , Junbo Wang , Wei Wang , Liang Wang , Tieniu Tan

E-commerce image search often takes a cropped image as the query, while each candidate is represented by full item images and structured text. This image-to-multimodal retrieval setting presents two asymmetries: a modality disparity -- a…

Information Retrieval · Computer Science 2026-05-19 Xinyu Sun , Huangyu Dai , Lingtao Mao , Zexin Zheng , Zihan Liang , Ben Chen , Chenyi Lei , Wenwu Ou

Text-to-image person re-identification (TIReID) aims to retrieve the target person from an image gallery via a textual description query. Recently, pre-trained vision-language models like CLIP have attracted significant attention and have…

Computer Vision and Pattern Recognition · Computer Science 2024-01-05 Weihao Li , Lei Tan , Pingyang Dai , Yan Zhang

The Visual Language Model, known for its robust cross-modal capabilities, has been extensively applied in various computer vision tasks. In this paper, we explore the use of CLIP (Contrastive Language-Image Pretraining), a vision-language…

Computer Vision and Pattern Recognition · Computer Science 2025-02-12 Huazhong Zhao , Lei Qi , Xin Geng

Text-Based Person Search (TBPS) aims to retrieve target person images from a large-scale gallery using natural language descriptions, posing fundamental challenges in cross-modal representation learning. Existing methods often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Jing Liu , Donglai Wei , Yang Liu , Sipeng Zhang , Tong Yang , Wei Zhou , Weiping Ding , Victor C. M. Leung

Remote sensing cross-modal text-image retrieval (RSCTIR) has gained attention for its utility in information mining. However, challenges remain in effectively integrating global and local information due to variations in remote sensing…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Zengbao Sun , Ming Zhao , Gaorui Liu , André Kaup

Aligning features from different modalities, is one of the most fundamental challenges for cross-modal tasks. Although pre-trained vision-language models can achieve a general alignment between image and text, they often require…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Ziqi Jiang , Yanghao Wang , Long Chen

Text-image cross-modal retrieval is a challenging task in the field of language and vision. Most previous approaches independently embed images and sentences into a joint embedding space and compare their similarities. However, previous…

Computer Vision and Pattern Recognition · Computer Science 2019-09-13 Zihao Wang , Xihui Liu , Hongsheng Li , Lu Sheng , Junjie Yan , Xiaogang Wang , Jing Shao

Vision-language models like CLIP have shown impressive capabilities in aligning images and text, but they often struggle with lengthy and detailed text descriptions because of their training focus on short and concise captions. We present…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Hyungyu Choi , Young Kyun Jang , Chanho Eom

Visible-infrared cross-modality person re-identification is a challenging ReID task, which aims to retrieve and match the same identity's images between the heterogeneous visible and infrared modalities. Thus, the core of this task is to…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Tengfei Liang , Yi Jin , Yajun Gao , Wu Liu , Songhe Feng , Tao Wang , Yidong Li

Image fusion aims to synthesize a single high-quality image from a pair of inputs captured under challenging conditions, such as differing exposure levels or focal depths. A core challenge lies in effectively handling disparities in dynamic…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Mingwei Tang , Jiahao Nie , Guang Yang , Ziqing Cui , Jie Li

Visible-Infrared Person Re-Identification (VI-ReID) is a challenging task due to the large modality discrepancy between visible and infrared images, which complicates the alignment of their features into a suitable common space. Moreover,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Neng Dong , Shuanglin Yan , Liyan Zhang , Jinhui Tang

In the realm of Text-Based Person Search (TBPS), mainstream methods aim to explore more efficient interaction frameworks between text descriptions and visual data. However, recent approaches encounter two principal challenges. Firstly, the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Lei Tan , Weihao Li , Pingyang Dai , Jie Chen , Liujuan Cao , Rongrong Ji