中文
相关论文

相关论文: FitPro: A Zero-Shot Framework for Interactive Text…

200 篇论文

Existing methods for interactive image retrieval have demonstrated the merit of integrating user feedback, improving retrieval results. However, most current systems rely on restricted forms of user feedback, such as binary relevance…

计算机视觉与模式识别 · 计算机科学 2018-12-24 Xiaoxiao Guo , Hui Wu , Yu Cheng , Steven Rennie , Gerald Tesauro , Rogerio Schmidt Feris

In the past few years, cross-modal image-text retrieval (ITR) has experienced increased interest in the research community due to its excellent research value and broad real-world application. It is designed for the scenarios where the…

信息检索 · 计算机科学 2022-11-21 Min Cao , Shiping Li , Juntao Li , Liqiang Nie , Min Zhang

Text-based person retrieval aims to identify a target individual from an image gallery using a natural language description. Existing methods primarily focus on appearance-driven cross-modal retrieval, yet face significant challenges due to…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yingjia Xu , Jinlin Wu , Daming Gao , Zhen Chen , Yang Yang , Min Cao , Mang Ye , Zhen Lei

This work introduces Text-based Aerial-Ground Person Retrieval (TAG-PR), which aims to retrieve person images from heterogeneous aerial and ground views with textual descriptions. Unlike traditional Text-based Person Retrieval (T-PR), which…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Xinyu Zhou , Yu Wu , Jiayao Ma , Wenhao Wang , Min Cao , Mang Ye

Pedestrian attribute recognition (PAR) aims to predict the attributes of a target pedestrian in a surveillance system. Existing methods address the PAR problem by training a multi-label classifier with predefined attribute classes. However,…

计算机视觉与模式识别 · 计算机科学 2023-11-27 Yue Zhang , Suchen Wang , Shichao Kan , Zhenyu Weng , Yigang Cen , Yap-peng Tan

We present T-Rex2, a highly practical model for open-set object detection. Previous open-set object detection methods relying on text prompts effectively encapsulate the abstract concept of common objects, but struggle with rare or complex…

计算机视觉与模式识别 · 计算机科学 2024-03-22 Qing Jiang , Feng Li , Zhaoyang Zeng , Tianhe Ren , Shilong Liu , Lei Zhang

Composed Image Retrieval (CIR) aims to retrieve a target image based on a reference image and conditioning text, enabling controllable image searches. The mainstream Zero-Shot (ZS) CIR methods bypass the need for expensive training CIR…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Jaeseok Byun , Seokhyeon Jeong , Wonjae Kim , Sanghyuk Chun , Taesup Moon

Text-based person search (TBPS) aims to retrieve the images of the target person from a large image gallery based on a given natural language description. Existing methods are dominated by training models with parallel image-text pairs,…

计算机视觉与模式识别 · 计算机科学 2023-08-07 Yang Bai , Jingyao Wang , Min Cao , Chen Chen , Ziqiang Cao , Liqiang Nie , Min Zhang

Person retrieval has attracted rising attention. Existing methods are mainly divided into two retrieval modes, namely image-only and text-only. However, they are unable to make full use of the available information and are difficult to meet…

计算机视觉与模式识别 · 计算机科学 2025-05-21 Delong Liu , Haiwen Li , Zhaohui Hou , Zhicheng Zhao , Fei Su , Yuan Dong

Text image super-resolution (Text-SR) requires more than visually plausible detail synthesis: slight errors in stroke topology may alter character identity and break readability. Existing methods improve text fidelity with stronger…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Zihang Xu , Xiaoyang Liu , Zheng Chen , Yulun Zhang , Xiaokang Yang

Scene text retrieval aims to find all images containing the query text from an image gallery. Current efforts tend to adopt an Optical Character Recognition (OCR) pipeline, which requires complicated text detection and/or recognition…

计算机视觉与模式识别 · 计算机科学 2024-08-02 Gangyan Zeng , Yuan Zhang , Jin Wei , Dongbao Yang , Peng Zhang , Yiwen Gao , Xugong Qin , Yu Zhou

Text-to-image person re-identification (TI-ReID) relies on natural-language text description to retrieve top matching individuals from a large gallery of images. While recent large vision-language models (VLMs) achieve strong retrieval…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Shakeeb Murtaza , Aryan Shukla , Rajarshi Bhattacharya , Maguelonne Heritier , Eric Granger

In this paper, we present TMR, a simple yet effective approach for text to 3D human motion retrieval. While previous work has only treated retrieval as a proxy evaluation metric, we tackle it as a standalone task. Our method extends the…

计算机视觉与模式识别 · 计算机科学 2023-08-28 Mathis Petrovich , Michael J. Black , Gül Varol

The development of image time series retrieval (ITSR) methods is a growing research interest in remote sensing (RS). Given a user-defined image time series (i.e., the query time series), ITSR methods search and retrieve from large archives…

计算机视觉与模式识别 · 计算机科学 2025-07-16 Genc Hoxha , Olivér Angyal , Begüm Demir

Human action recognition is pivotal in computer vision, with applications ranging from surveillance to human-robot interaction. Despite the effectiveness of supervised skeleton-based methods, their reliance on exhaustive annotation limits…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Yuxi Zhou , Zhengbo Zhang , Jingyu Pan , Zhiyu Lin , Zhigang Tu

Pedestrian intention prediction is essential for autonomous driving in complex urban environments. Conventional approaches depend on supervised learning over frame sequences and require extensive retraining to adapt to new scenarios. Here,…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Pallavi Zambare , Venkata Nikhil Thanikella , Ying Liu

Recent advances in zero-shot text-to-speech (TTS) have enabled accurate imitation of reference speech in terms of both speaking style and speaker timbre. However, achieving disentangled control over these aspects from separate references…

音频与语音处理 · 电气工程与系统科学 2026-05-26 Yoonhyung Lee , Hyunsin Park , Jinhwan Park , Jinkyu Lee

Zero-shot composed image retrieval (ZS-CIR) is a rapidly growing area with significant practical applications, allowing users to retrieve a target image by providing a reference image and a relative caption describing the desired…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Yongcong Ye , Kai Zhang , Yanghai Zhang , Enhong Chen , Longfei Li , Jun Zhou

Zero-shot scene understanding in real-world settings presents major challenges due to the complexity and variability of natural scenes, where models must recognize new objects, actions, and contexts without prior labeled examples. This work…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Manjunath Prasad Holenarasipura Rajiv , B. M. Vidyavathi

Detection of pedestrians on embedded devices, such as those on-board of robots and drones, has many applications including road intersection monitoring, security, crowd monitoring and surveillance, to name a few. However, the problem can be…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Mohamed Afifi , Yara Ali , Karim Amer , Mahmoud Shaker , Mohamed Elhelw