English
Related papers

Related papers: FitPro: A Zero-Shot Framework for Interactive Text…

200 papers

Text-based pedestrian search (TBPS) in full images aims to locate a target pedestrian in untrimmed images using natural language descriptions. However, in complex scenes with multiple pedestrians, existing methods are limited by…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Zengli Luo , Canlong Zhang , Zhixin Li , Zhiwen Wang , Chunrong Wei

Text-to-Image Person Retrieval (TIPR) is a cross-modal matching task designed to identify the person images that best correspond to a given textual description. The key difficulty in TIPR is to realize robust correspondence between the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Hao Yin , Xin Man , Feiyu Chen , Jie Shao , Heng Tao Shen

Finding target persons in full scene images with a query of text description has important practical applications in intelligent video surveillance.However, different from the real-world scenarios where the bounding boxes are not available,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Shizhou Zhang , De Cheng , Wenlong Luo , Yinghui Xing , Duo Long , Hao Li , Kai Niu , Guoqiang Liang , Yanning Zhang

Most existing methods for text-based person retrieval focus on text-to-image person retrieval. Nevertheless, due to the lack of dynamic information provided by isolated frames, the performance is hampered when the person is obscured or…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Xu Zhang , Fan Ni , Guan-Nan Dong , Aichun Zhu , Jianhui Wu , Mingcheng Ni , Hui Liu

Text-Pedestrian Image Retrieval aims to use the text describing pedestrian appearance to retrieve the corresponding pedestrian image. This task involves not only modality discrepancy, but also the challenge of the textual diversity of…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Huafeng Li , Shedan Yang , Yafei Zhang , Dapeng Tao , Zhengtao Yu

Text-based person retrieval (TPR) is a challenging task that involves retrieving a specific individual based on a textual description. Despite considerable efforts to bridge the gap between vision and language, the significant differences…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Yiwei Ma , Xiaoshuai Sun , Jiayi Ji , Guannan Jiang , Weilin Zhuang , Rongrong Ji

Visual Place Recognition (VPR) has evolved from handcrafted descriptors to deep learning approaches, yet significant challenges remain. Current approaches, including Vision Foundation Models (VFMs) and Multimodal Large Language Models…

Machine Learning · Computer Science 2025-09-03 Jintao Cheng , Weibin Li , Jiehao Luo , Xiaoyu Tang , Zhijian He , Jin Wu , Yao Zou , Wei Zhang

The goal of Text-to-Image Person Retrieval (TIPR) is to retrieve specific person images according to the given textual descriptions. A primary challenge in this task is bridging the substantial representational gap between visual and…

Computation and Language · Computer Science 2025-01-20 Delong Liu , Haiwen Li , Zhicheng Zhao , Yuan Dong

We propose GOTPR, a robust place recognition method designed for outdoor environments where GPS signals are unavailable. Unlike existing approaches that use point cloud maps, which are large and difficult to store, GOTPR leverages scene…

Robotics · Computer Science 2025-05-23 Donghwi Jung , Keonwoo Kim , Seong-Woo Kim

Text-based person search aims to retrieve the matched pedestrians from a large-scale image database according to the text description. The core difficulty of this task is how to extract effective details from pedestrian images and texts,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Wei Shen , Ming Fang , Yuxia Wang , Jiafeng Xiao , Diping Li , Huangqun Chen , Ling Xu , Weifeng Zhang

As a fundamental and challenging task in bridging language and vision domains, Image-Text Retrieval (ITR) aims at searching for the target instances that are semantically relevant to the given query from the other modality, and its key…

Computer Vision and Pattern Recognition · Computer Science 2023-01-18 Yan Zhang , Zhong Ji , Di Wang , Yanwei Pang , Xuelong Li

Person text-image matching, also known as text based person search, aims to retrieve images of specific pedestrians using text descriptions. Although person text-image matching has made great research progress, existing methods still face…

Computer Vision and Pattern Recognition · Computer Science 2022-11-22 Fan Li , Hang Zhou , Huafeng Li , Yafei Zhang , Zhengtao Yu

Text-motion retrieval aims to learn a semantically aligned latent space between natural language descriptions and 3D human motion skeleton sequences, enabling bidirectional search across the two modalities. Most existing methods use a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Yao Zhang , Zhuchenyang Liu , Yanlan He , Thomas Ploetz , Yu Xiao

Visible-infrared person re-identification (VI-ReID) enables cross-modality identity matching for all-day surveillance, yet existing methods predominantly focus on the image level or rely heavily on costly identity annotations. While…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Zhiyong Li , Wei Jiang , Haojie Liu , Mingyu Wang , Wanchong Xu , Weijie Mao

Text-Based Person Search (TBPS) aims to retrieve pedestrian images from large galleries using natural language descriptions. This task, essential for public safety applications, is hindered by cross-modal discrepancies and ambiguous user…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Zequn Xie

Monocular 3D object detection typically relies on pseudo-labeling techniques to reduce dependency on real-world annotations. Recent advances demonstrate that deterministic linguistic cues can serve as effective auxiliary weak supervision…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Chupeng Liu , Jiyong Rao , Shangquan Sun , Runkai Zhao , Weidong Cai

In the field of Image-Text Retrieval (ITR), recent advancements have leveraged large-scale Vision-Language Pretraining (VLP) for Fine-Grained (FG) instance-level retrieval, achieving high accuracy at the cost of increased computational…

Information Retrieval · Computer Science 2026-01-19 Mikel Williams-Lekuona , Georgina Cosma

Zero-shot sketch-based image retrieval (ZS-SBIR) is a specific cross-modal retrieval task for searching natural images given free-hand sketches under the zero-shot scenario. Most existing methods solve this problem by simultaneously…

Computer Vision and Pattern Recognition · Computer Science 2022-05-09 Xinxun Xu , Muli Yang , Yanhua Yang , Hao Wang

Federated Magnetic Resonance Imaging (MRI) reconstruction enables multiple hospitals to collaborate distributedly without aggregating local data, thereby protecting patient privacy. However, the data heterogeneity caused by different MRI…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Chun-Mei Feng , Bangjun Li , Xinxing Xu , Yong Liu , Huazhu Fu , Wangmeng Zuo

Large-scale pretrained image-text models have shown incredible zero-shot performance in a handful of tasks, including video ones such as action recognition and text-to-video retrieval. However, these models have not been adapted to video,…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Santiago Castro , Fabian Caba Heilbron
‹ Prev 1 2 3 10 Next ›