中文
相关论文

相关论文: Find Them All: Unveiling MLLMs for Versatile Perso…

200 篇论文

The growing importance of person reidentification in computer vision has highlighted the need for more extensive and diverse datasets. In response, we introduce the ENTIRe-ID dataset, an extensive collection comprising over 4.45 million…

计算机视觉与模式识别 · 计算机科学 2024-06-03 Serdar Yildiz , Ahmet Nezih Kasim

Person re-identification (ReId), a crucial task in surveillance, involves matching individuals across different camera views. The advent of Deep Learning, especially supervised techniques like Convolutional Neural Networks and Attention…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Andrea Asperti , Salvatore Fiorilla , Simone Nardi , Lorenzo Orsini

The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LVLMs have begun to address this need. However, their…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Fanqing Meng , Jin Wang , Chuanhao Li , Quanfeng Lu , Hao Tian , Jiaqi Liao , Xizhou Zhu , Jifeng Dai , Yu Qiao , Ping Luo , Kaipeng Zhang , Wenqi Shao

Most existing person re-identification (ReID) methods rely only on the spatial appearance information from either one or multiple person images, whilst ignore the space-time cues readily available in video or image-sequence data. Moreover,…

计算机视觉与模式识别 · 计算机科学 2016-11-29 Xiaolong Ma , Xiatian Zhu , Shaogang Gong , Xudong Xie , Jianming Hu , Kin-Man Lam , Yisheng Zhong

With the rapid advancement of Multi-modal Large Language Models (MLLMs), their capability in understanding both images and text has greatly improved. However, their potential for leveraging multi-modal contextual information in…

人工智能 · 计算机科学 2025-08-08 Zhenghao Liu , Xingsheng Zhu , Tianshuo Zhou , Xinyi Zhang , Xiaoyuan Yi , Yukun Yan , Ge Yu , Maosong Sun

Medical image re-identification (MedReID) is under-explored so far, despite its critical applications in personalized healthcare and privacy protection. In this paper, we introduce a thorough benchmark and a unified model for this problem.…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Yuan Tian , Kaiyuan Ji , Rongzhao Zhang , Yankai Jiang , Chunyi Li , Xiaosong Wang , Guangtao Zhai

Individuals with visual impairments, encompassing both partial and total difficulties in visual perception, are referred to as visually impaired (VI) people. An estimated 2.2 billion individuals worldwide are affected by visual impairments.…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Bufang Yang , Lixing He , Kaiwei Liu , Zhenyu Yan

Learning to re-identify or retrieve a group of people across non-overlapped camera systems has important applications in video surveillance. However, most existing methods focus on (single) person re-identification (re-id), ignoring the…

计算机视觉与模式识别 · 计算机科学 2021-04-30 Yichao Yan , Jie Qin , Bingbing Ni , Jiaxin Chen , Li Liu , Fan Zhu , Wei-Shi Zheng , Xiaokang Yang , Ling Shao

Person re-identification (Re-ID) aims to match person images across different camera views, with occluded Re-ID addressing scenarios where pedestrians are partially visible. While pre-trained vision-language models have shown effectiveness…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Rui Zhi , Zhen Yang , Haiyang Zhang

Text-to-image person re-identification (TIReID) retrieves pedestrian images of the same identity based on a query text. However, existing methods for TIReID typically treat it as a one-to-one image-text matching problem, only focusing on…

计算机视觉与模式识别 · 计算机科学 2023-10-18 Shuanglin Yan , Neng Dong , Jun Liu , Liyan Zhang , Jinhui Tang

Visual-language pre-training has achieved remarkable success in many multi-modal tasks, largely attributed to the availability of large-scale image-text datasets. In this work, we demonstrate that Multi-modal Large Language Models (MLLMs)…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Yanqing Liu , Kai Wang , Wenqi Shao , Ping Luo , Yu Qiao , Mike Zheng Shou , Kaipeng Zhang , Yang You

The development of large language models (LLMs) has significantly enhanced the capabilities of multimodal LLMs (MLLMs) as general assistants. However, lack of user-specific knowledge still restricts their application in human's daily life.…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Haoran Hao , Jiaming Han , Changsheng Li , Yu-Feng Li , Xiangyu Yue

Multimodal Large Language Models (MLLMs) have advanced VQA and now support Vision-DeepResearch systems that use search engines for complex visual-textual fact-finding. However, evaluating these visual and textual search abilities is still…

With the continuous advancement of large language models (LLMs), it is essential to create new benchmarks to effectively evaluate their expanding capabilities and identify areas for improvement. This work focuses on multi-image reasoning,…

Recent researchers have proposed using event cameras for person re-identification (ReID) due to their promising performance and better balance in terms of privacy protection, event camera-based person ReID has attracted significant…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Xiao Wang , Qian Zhu , Shujuan Wu , Bo Jiang , Shiliang Zhang

Visible-infrared person re-identification (VI-ReID) is a challenging and essential task, which aims to retrieve a set of person images over visible and infrared camera views. In order to mitigate the impact of large modality discrepancy…

计算机视觉与模式识别 · 计算机科学 2023-07-26 Haojie Liu , Daoxun Xia , Wei Jiang , Chao Xu

Visible-Infrared Person Re-Identification (VI-ReID) is a challenging retrieval task under complex modality changes. Existing methods usually focus on extracting discriminative visual features while ignoring the reliability and commonality…

计算机视觉与模式识别 · 计算机科学 2022-12-02 Hu Lu , Xuezhang Zou , Pingping Zhang

Person re-identification (Re-ID) aims to match a target person across camera views at different locations and times. Existing Re-ID studies focus on the short-term cloth-consistent setting, under which a person re-appears in different…

计算机视觉与模式识别 · 计算机科学 2020-10-08 Xuelin Qian , Wenxuan Wang , Li Zhang , Fangrui Zhu , Yanwei Fu , Tao Xiang , Yu-Gang Jiang , Xiangyang Xue

Person re-identification (re-id) is the task of recognizing and matching persons at different locations recorded by cameras with non-overlapping views. One of the main challenges of re-id is the large variance in person poses and camera…

计算机视觉与模式识别 · 计算机科学 2018-03-26 Andreas Eberle

In text-to-image person retrieval tasks, the diversity of natural language expressions and the implicitness of visual semantics often lead to the problem of Expression Drift, where semantically equivalent texts exhibit significant feature…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Chao Yuan , Yujian Zhao , Haoxuan Xu , Guanglin Niu