English
Related papers

Related papers: Multilingual Text-to-Image Person Retrieval via Bi…

200 papers

In this paper we propose a model to learn multimodal multilingual representations for matching images and sentences in different languages, with the aim of advancing multilingual versions of image search and image understanding. Our model…

Computation and Language · Computer Science 2017-07-25 Spandana Gella , Rico Sennrich , Frank Keller , Mirella Lapata

Human intelligence can retrieve any person according to both visual and language descriptions. However, the current computer vision community studies specific person re-identification (ReID) tasks in different scenarios separately, which…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Weizhen He , Yiheng Deng , Yunfeng Yan , Feng Zhu , Yizhou Wang , Lei Bai , Qingsong Xie , Donglian Qi , Wanli Ouyang , Shixiang Tang

Text-based person search aims to retrieve the corresponding person images in an image database by virtue of a describing sentence about the person, which poses great potential for various applications such as video surveillance. Extracting…

Computer Vision and Pattern Recognition · Computer Science 2019-11-28 Ya Jing , Chenyang Si , Junbo Wang , Wei Wang , Liang Wang , Tieniu Tan

Existing mainstream approaches follow the encoder-decoder paradigm for generating radiology reports. They focus on improving the network structure of encoders and decoders, which leads to two shortcomings: overlooking the modality gap and…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Yuanjiang Luo , Hongxiang Li , Xuan Wu , Meng Cao , Xiaoshuang Huang , Zhihong Zhu , Peixi Liao , Hu Chen , Yi Zhang

Embodied agents have achieved prominent performance in following human instructions to complete tasks. However, the potential of providing instructions informed by texts and images to assist humans in completing tasks remains underexplored.…

Computation and Language · Computer Science 2023-05-04 Yujie Lu , Pan Lu , Zhiyu Chen , Wanrong Zhu , Xin Eric Wang , William Yang Wang

Unified multimodal understanding and generation models recently have achieve significant improvement in image generation capability, yet a large gap remains in instruction following and detail preservation compared to systems that tightly…

We propose a visual-linguistic representation learning approach within a self-supervised learning framework by introducing a new operation, loss, and data augmentation strategy. First, we generate diverse features for the image-text…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Jaeyoo Park , Bohyung Han

Audio-text retrieval (ATR), which retrieves a relevant caption given an audio clip (A2T) and vice versa (T2A), has recently attracted much research attention. Existing methods typically aggregate information from each modality into a single…

Sound · Computer Science 2024-03-18 Qian Wang , Jia-Chen Gu , Zhen-Hua Ling

Cross-lingual topic modeling aims to uncover shared semantic themes across languages. Several methods have been proposed to address this problem, leveraging both traditional and neural approaches. While previous methods have achieved some…

Computation and Language · Computer Science 2025-10-06 Tien Phat Nguyen , Vu Minh Ngo , Tung Nguyen , Linh Van Ngo , Duc Anh Nguyen , Sang Dinh , Trung Le

Real-world infrared imagery presents unique challenges for vision-language models due to the scarcity of aligned text data and domain-specific characteristics. Although existing methods have advanced the field, their reliance on synthetic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Zhe Cao , Jin Zhang , Ruiheng Zhang

RGB-Infrared (IR) person re-identification is very challenging due to the large cross-modality variations between RGB and IR images. The key solution is to learn aligned features to the bridge RGB and IR modalities. However, due to the lack…

Computer Vision and Pattern Recognition · Computer Science 2020-02-19 Guan-An Wang , Tianzhu Zhang. Yang Yang , Jian Cheng , Jianlong Chang , Xu Liang , Zengguang Hou

Region-based image retrieval (RBIR) technique is revisited. In early attempts at RBIR in the late 90s, researchers found many ways to specify region-based queries and spatial relationships; however, the way to characterize the regions, such…

Multimedia · Computer Science 2017-09-27 Ryota Hinami , Yusuke Matsui , Shin'ichi Satoh

Multimodal learning is a recent challenge that extends unimodal learning by generalizing its domain to diverse modalities, such as texts, images, or speech. This extension requires models to process and relate information from multiple…

Information Retrieval · Computer Science 2022-09-29 Cheng-An Hsieh , Cheng-Ping Hsieh , Pu-Jen Cheng

Text-based person search aims to retrieve images of a certain pedestrian by a textual description. The key challenge of this task is to eliminate the inter-modality gap and achieve the feature alignment across modalities. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2021-12-14 Shiping Li , Min Cao , Min Zhang

Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve target images given a compositional query, consisting of a reference image and a modifying text-without relying on annotated training data. Existing approaches often generate a…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Rong-Cheng Tu , Wenhao Sun , Hanzhe You , Yingjie Wang , Jiaxing Huang , Li Shen , Dacheng Tao

Text image translation (TIT) aims to translate the source texts embedded in the image to target translations, which has a wide range of applications and thus has important research value. However, current studies on TIT are confronted with…

Computation and Language · Computer Science 2023-06-05 Zhibin Lan , Jiawei Yu , Xiang Li , Wen Zhang , Jian Luan , Bin Wang , Degen Huang , Jinsong Su

Text-to-image person re-identification (TIReID) aims to retrieve person images from a large gallery given free-form textual descriptions. TIReID is challenging due to the substantial modality gap between visual appearances and textual…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Giyeol Kim , Chanho Eom

How humans can effectively and efficiently acquire images has always been a perennial question. A classic solution is text-to-image retrieval from an existing database; however, the limited database typically lacks creativity. By contrast,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Leigang Qu , Haochuan Li , Tan Wang , Wenjie Wang , Yongqi Li , Liqiang Nie , Tat-Seng Chua

Text-based person retrieval aims to identify specific individuals within an image database using textual descriptions. Due to the high cost of annotation and privacy protection, researchers resort to synthesized data for the paradigm of…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Hang Yu , Jiahao Wen , Zhedong Zheng

RGB-Infrared (IR) person re-identification aims to retrieve person-of-interest from heterogeneous cameras, easily suffering from large image modality discrepancy caused by different sensing wavelength ranges. Existing work usually minimizes…

Computer Vision and Pattern Recognition · Computer Science 2021-07-27 Lin Wan , Zongyuan Sun , Qianyan Jing , Yehansen Chen , Lijing Lu , Zhihang Li
‹ Prev 1 4 5 6 7 8 10 Next ›