English
Related papers

Related papers: TIGER-FG: Text-Guided Implicit Fine-Grained Ground…

200 papers

Image-text matching is gaining a leading role among tasks involving the joint understanding of vision and language. In literature, this task is often used as a pre-training objective to forge architectures able to jointly deal with images…

Computer Vision and Pattern Recognition · Computer Science 2022-08-01 Nicola Messina , Matteo Stefanini , Marcella Cornia , Lorenzo Baraldi , Fabrizio Falchi , Giuseppe Amato , Rita Cucchiara

Discriminative representation is essential to keep a unique identifier for each target in Multiple object tracking (MOT). Some recent MOT methods extract features of the bounding box region or the center point as identity embeddings.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-09 Hao Ren , Shoudong Han , Huilin Ding , Ziwen Zhang , Hongwei Wang , Faquan Wang

State-of-the-art face recognition (FR) models often experience a significant performance drop when dealing with facial images in surveillance scenarios where images are in low quality and often corrupted with noise. Leveraging facial…

Computer Vision and Pattern Recognition · Computer Science 2023-12-18 Md Mahedi Hasan , Shoaib Meraj Sami , Nasser Nasrabadi

Fine-grained visual categorization is to recognize hundreds of subcategories belonging to the same basic-level category, which is a highly challenging task due to the quite subtle and local visual distinctions among similar subcategories.…

Computer Vision and Pattern Recognition · Computer Science 2019-02-21 Xiangteng He , Yuxin Peng

Recent studies extend the autoregression paradigm to text-to-image generation, achieving performance comparable to diffusion models. However, our new PairComp benchmark -- featuring test cases of paired prompts with similar syntax but…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Kaihang Pan , Wendong Bu , Yuruo Wu , Yang Wu , Kai Shen , Yunfei Li , Hang Zhao , Juncheng Li , Siliang Tang , Yueting Zhuang

Text-to-image (T2I) generation has achieved remarkable progress in instruction following and aesthetics. However, a persistent challenge is the prevalence of physical artifacts, such as anatomical and structural flaws, which severely…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Jia Wang , Jie Hu , Xiaoqi Ma , Hanghang Ma , Yanbing Zeng , Xiaoming Wei

Scene Graph Generation (SGG) aims to structurally and comprehensively represent objects and their connections in images, it can significantly benefit scene understanding and other related downstream tasks. Existing SGG models often struggle…

Computer Vision and Pattern Recognition · Computer Science 2023-06-26 Qianji Di , Wenxi Ma , Zhongang Qi , Tianxiang Hou , Ying Shan , Hanzi Wang

Fine-grained image search is still a challenging problem due to the difficulty in capturing subtle differences regardless of pose variations of objects from fine-grained categories. In practice, a dynamic inventory with new fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2018-07-09 Kevin Lin , Fan Yang , Qiaosong Wang , Robinson Piramuthu

Text-based image retrieval has seen considerable progress in recent years. However, the performance of existing methods suffers in real life since the user is likely to provide an incomplete description of an image, which often leads to…

Computer Vision and Pattern Recognition · Computer Science 2021-08-12 Guanyu Cai , Jun Zhang , Xinyang Jiang , Yifei Gong , Lianghua He , Fufu Yu , Pai Peng , Xiaowei Guo , Feiyue Huang , Xing Sun

Fashion image retrieval task aims to search relevant clothing items of a query image from the gallery. The previous recipes focus on designing different distance-based loss functions, pulling relevant pairs to be close and pushing…

Computer Vision and Pattern Recognition · Computer Science 2023-03-09 Jinkuan Zhu , Hao Huang , Qiao Deng , Xiyao Li

Recent CLIP-based few-shot semantic segmentation methods introduce class-level textual priors to assist segmentation by typically using a single prompt (e.g., a photo of class). However, these approaches often result in incomplete…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Qiang Jiao , Bin Yan , Yi Yang , Mengrui Shi , Qiang Zhang

Current multimodal information retrieval studies mainly focus on single-image inputs, which limits real-world applications involving multiple images and text-image interleaved content. In this work, we introduce the text-image interleaved…

Computation and Language · Computer Science 2025-02-19 Xin Zhang , Ziqi Dai , Yongqi Li , Yanzhao Zhang , Dingkun Long , Pengjun Xie , Meishan Zhang , Jun Yu , Wenjie Li , Min Zhang

While generative modeling has become prevalent across numerous research fields, its integration into the realm of image retrieval remains largely unexplored and underjustified. In this paper, we present a novel methodology, reframing image…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Yidan Zhang , Ting Zhang , Dong Chen , Yujing Wang , Qi Chen , Xing Xie , Hao Sun , Weiwei Deng , Qi Zhang , Fan Yang , Mao Yang , Qingmin Liao , Jingdong Wang , Baining Guo

Industrial-scale recommender systems rely on a cascade pipeline in which the retrieval stage must return a high-recall candidate set from billions of items under tight latency. Existing solutions either (i) suffer from limited…

Information Retrieval · Computer Science 2026-04-02 Yijia Sun , Shanshan Huang , Zhiyuan Guan , Qiang Luo , Ruiming Tang , Kun Gai , Guorui Zhou

Fine-grained visual classification can be addressed by deep representation learning under supervision of manually pre-defined targets (e.g., one-hot or the Hadamard codes). Such target coding schemes are less flexible to model inter-class…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Kangjun Liu , Ke Chen , Kui Jia

Most existing methods in vision-language retrieval match two modalities by either comparing their global feature vectors which misses sufficient information and lacks interpretability, detecting objects in images or videos and aligning the…

Computer Vision and Pattern Recognition · Computer Science 2022-10-04 Xiaohan Zou , Changqiao Wu , Lele Cheng , Zhongyuan Wang

Fine-grained recognition involves the classification of images from subordinate macro-categories, and it is challenging due to small inter-class differences. To overcome this, most methods perform discriminative feature selection enabled by…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Edwin Arkel Rios , Min-Chun Hu , Bo-Cheng Lai

Image and text retrieval is one of the foundational tasks in the vision and language domain with multiple real-world applications. State-of-the-art approaches, e.g. CLIP, ALIGN, represent images and texts as dense embeddings and calculate…

Computer Vision and Pattern Recognition · Computer Science 2023-02-09 Chen Chen , Bowen Zhang , Liangliang Cao , Jiguang Shen , Tom Gunter , Albin Madappally Jose , Alexander Toshev , Jonathon Shlens , Ruoming Pang , Yinfei Yang

Interactive Text-to-image retrieval (I-TIR) is an important enabler for a wide range of state-of-the-art services in domains such as e-commerce and education. However, current methods rely on finetuned Multimodal Large Language Models…

Information Retrieval · Computer Science 2025-07-11 Zijun Long , Kangheng Liang , Gerardo Aragon-Camarasa , Richard Mccreadie , Paul Henderson

Identifying user intent from mobile UI operation trajectories is critical for advancing UI understanding and enabling task automation agents. While Multimodal Large Language Models (MLLMs) excel at video understanding tasks, their real-time…

Artificial Intelligence · Computer Science 2025-12-23 Zhe Yang , Xiaoshuang Sheng , Zhengnan Zhang , Jidong Wu , Zexing Wang , Xin He , Shenghua Xu , Guanjing Xiong