English
Related papers

Related papers: MARS: Paying more attention to visual attributes f…

200 papers

Current metrics for text-to-image models typically rely on statistical metrics which inadequately represent the real preference of humans. Although recent work attempts to learn these preferences via human annotated images, they reduce the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Sixian Zhang , Bohan Wang , Junqiang Wu , Yan Li , Tingting Gao , Di Zhang , Zhongyuan Wang

Most image-text retrieval work adopts binary labels indicating whether a pair of image and text matches or not. Such a binary indicator covers only a limited subset of image-text semantic relations, which is insufficient to represent…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Zheng Li , Caili Guo , Zerun Feng , Jenq-Neng Hwang , Ying Jin , Yufeng Zhang

Person search aims to localize specific a target person from a gallery set of images with various scenes. As the scene of moving pedestrian changes, the captured person image inevitably bring in lots of background noise and foreground noise…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Yimin Jiang , Huibing Wang , Jinjia Peng , Xianping Fu , Yang Wang

Most existing text-to-image person retrieval methods usually assume that the training image-text pairs are perfectly aligned; however, the noisy correspondence(NC) issue (i.e., incorrect or unreliable alignment) exists due to poor image…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Runqing Zhang , Xue Zhou

Most existing text recognition methods are trained on large-scale synthetic datasets due to the scarcity of labeled real-world datasets. Synthetic images, however, cannot faithfully reproduce real-world scenarios, such as uneven…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Zhengmi Tang , Yuto Mitsui , Tomo Miyazaki , Shinichiro Omachi

Nowadays, we have witnessed the early progress on learning the association between voice and face automatically, which brings a new wave of studies to the computer vision community. However, most of the prior arts along this line (a) merely…

Computer Vision and Pattern Recognition · Computer Science 2021-03-15 Peisong Wen , Qianqian Xu , Yangbangyan Jiang , Zhiyong Yang , Yuan He , Qingming Huang

Text-to-image person re-identification (ReID) aims to retrieve images of a person based on a given textual description. The key challenge is to learn the relations between detailed information from visual and textual modalities. Existing…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Dixuan Lin , Yixing Peng , Jingke Meng , Wei-Shi Zheng

Person retrieval faces many challenges including cluttered background, appearance variations (e.g., illumination, pose, occlusion) among different camera views and the similarity among different person's images. To address these issues, we…

Computer Vision and Pattern Recognition · Computer Science 2019-04-19 Lei Qi , Jing Huo , Lei Wang , Yinghuan Shi , Yang Gao

Image-text matching (ITM) is a fundamental problem in computer vision. The key issue lies in jointly learning the visual and textual representation to estimate their similarity accurately. Most existing methods focus on feature enhancement…

Computer Vision and Pattern Recognition · Computer Science 2024-06-28 Xuri Ge , Fuhai Chen , Songpei Xu , Fuxiang Tao , Jie Wang , Joemon M. Jose

Diffusion-weighted magnetic resonance imaging (dMRI) is widely used to assess the brain white matter. One of the most common computations in dMRI involves cross-subject tract-specific analysis, whereby dMRI-derived biomarkers are compared…

Image and Video Processing · Electrical Eng. & Systems 2023-07-12 Davood Karimi , Hamza Kebiri , Ali Gholipour

Person re-identification (re-id) is the task of recognizing and matching persons at different locations recorded by cameras with non-overlapping views. One of the main challenges of re-id is the large variance in person poses and camera…

Computer Vision and Pattern Recognition · Computer Science 2018-03-26 Andreas Eberle

For a given scene, humans can easily reason for the locations and pose to place objects. Designing a computational model to reason about these affordances poses a significant challenge, mirroring the intuitive reasoning abilities of humans.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Rishubh Parihar , Harsh Gupta , Sachidanand VS , R. Venkatesh Babu

Person re-identification aims at establishing the identity of a pedestrian from a gallery that contains images of multiple people obtained from a multi-camera system. Many challenges such as occlusions, drastic lighting and pose variations…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Guodong Ding , Salman Khan , Zhenmin Tang , Fatih Porikli

The performance of traditional text-image person retrieval task is easily affected by lighting variations due to imaging limitations of visible spectrum sensors. In recent years, cross-modal information fusion has emerged as an effective…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Yifei Deng , Chenglong Li , Zhenyu Chen , Zihen Xu , Jin Tang

Text-to-image person re-identification (TIReID) is a compelling topic in the cross-modal community, which aims to retrieve the target person based on a textual query. Although numerous TIReID methods have been proposed and achieved…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Yang Qin , Yingke Chen , Dezhong Peng , Xi Peng , Joey Tianyi Zhou , Peng Hu

Large vision-language models (VLMs) often benefit from intermediate visual cues, either injected via external tools or generated as latent visual tokens during reasoning, but these mechanisms still overlook fine-grained visual evidence…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Shuoshuo Zhang , Yizhen Zhang , Jingjing Fu , Lei Song , Jiang Bian , Yujiu Yang , Rui Wang

Fine-tuning Multimodal Large Language Models (MLLMs) with parameter-efficient methods like Low-Rank Adaptation (LoRA) is crucial for task adaptation. However, imbalanced training dynamics across modalities often lead to suboptimal accuracy…

Machine Learning · Computer Science 2026-03-03 Minkyoung Cho , Insu Jang , Shuowei Jin , Zesen Zhao , Adityan Jothi , Ethem F. Can , Min-Hung Chen , Z. Morley Mao

Person re-identification (re-ID) is a highly challenging task due to large variations of pose, viewpoint, illumination, and occlusion. Deep metric learning provides a satisfactory solution to person re-ID by training a deep network under…

Computer Vision and Pattern Recognition · Computer Science 2018-07-31 Rui Yu , Zhiyong Dou , Song Bai , Zhaoxiang Zhang , Yongchao Xu , Xiang Bai

Tactics, Techniques and Procedures (TTPs) represent sophisticated attack patterns in the cybersecurity domain, described encyclopedically in textual knowledge bases. Identifying TTPs in cybersecurity writing, often called TTP mapping, is an…

Machine Learning · Computer Science 2025-07-28 Tu Nguyen , Nedim Šrndić , Alexander Neth

Text-Pedestrian Image Retrieval aims to use the text describing pedestrian appearance to retrieve the corresponding pedestrian image. This task involves not only modality discrepancy, but also the challenge of the textual diversity of…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Huafeng Li , Shedan Yang , Yafei Zhang , Dapeng Tao , Zhengtao Yu
‹ Prev 1 3 4 5 6 7 10 Next ›