中文
相关论文

相关论文: MambaPro: Multi-Modal Object Re-Identification wit…

200 篇论文

With the rapid advancements in deep learning technologies, person re-identification (ReID) has witnessed remarkable performance improvements. However, the majority of prior works have traditionally focused on solving the problem via…

计算机视觉与模式识别 · 计算机科学 2024-01-15 Huiyuan Fu , Kuilong Cui , Chuanming Wang , Mengshi Qi , Huadong Ma

Person reidentification (ReID) is a very hot research topic in machine learning and computer vision, and many person ReID approaches have been proposed; however, most of these methods assume that the same person has the same clothes within…

计算机视觉与模式识别 · 计算机科学 2021-08-11 Zan Gao , Hongwei Wei , Weili Guan , Weizhi Nie , Meng Liu , Meng Wang

Current research on Multimodal Retrieval-Augmented Generation (MRAG) enables diverse multimodal inputs but remains limited to single-modality outputs, restricting expressive capacity and practical utility. In contrast, real-world…

信息检索 · 计算机科学 2025-08-11 Zhiyou Xiao , Qinhan Yu , Binghui Li , Geng Chen , Chong Chen , Wentao Zhang

The aim of multiple object tracking (MOT) is to detect all objects in a video and bind them into multiple trajectories. Generally, this process is carried out in two steps: detecting objects and associating them across frames based on…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Ruopeng Gao , Yuyao Wang , Chunxu Liu , Limin Wang

Text-to-image person re-identification (ReID) aims to search for pedestrian images of an interested identity via textual descriptions. It is challenging due to both rich intra-modal variations and significant inter-modal gaps. Existing…

计算机视觉与模式识别 · 计算机科学 2022-07-19 Zhiyin Shao , Xinyu Zhang , Meng Fang , Zhifeng Lin , Jian Wang , Changxing Ding

Vision Language Place Recognition (VLVPR) enhances robot localization performance by incorporating natural language descriptions from images. By utilizing language information, VLVPR directs robot place matching, overcoming the constraint…

计算机视觉与模式识别 · 计算机科学 2025-02-24 Tianyi Shang , Zhenyu Li , Pengjie Xu , Jinwei Qiao

Action quality assessment (AQA) aims to automatically quantify the execution quality of human actions in videos and is valuable for applications such as competitive sports judging. In multimodal AQA, quality evidence from different…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Qiqi Li , Pengfei Wang , Nenggan Zheng

Recent State Space Models (SSM), especially Mamba, have demonstrated impressive performance in visual modeling and possess superior model efficiency. However, the application of Mamba to visual tasks suffers inferior performance due to…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Fei Xie , Jiahao Nie , Yujin Tang , Wenkang Zhang , Hongshen Zhao

Object re-identification is of increasing importance in visual surveillance. Most existing works focus on re-identify individual from multiple cameras while the application of group re-identification (Re-ID) is rarely discussed. We redefine…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Hao Xiao

Re-Identification (ReID) is a critical technology in intelligent perception systems, especially within autonomous driving, where onboard cameras must identify pedestrians across views and time in real-time to support safe navigation and…

计算机视觉与模式识别 · 计算机科学 2025-08-18 Jialin Li , Shuqi Wu , Ning Wang

New retrieval tasks have always been emerging, thus urging the development of new retrieval models. However, instantiating a retrieval model for each new retrieval task is resource-intensive and time-consuming, especially for a retrieval…

信息检索 · 计算机科学 2023-03-24 Juhao Liang , Chen Zhang , Zhengyang Tang , Jie Fu , Dawei Song , Benyou Wang

"This work has been submitted to the lEEE for possible publication. Copyright may be transferred without noticeafter which this version may no longer be accessible." Time series modeling serves as the cornerstone of real-world applications,…

机器学习 · 计算机科学 2025-04-04 Sijie Xiong , Shuqing Liu , Cheng Tang , Fumiya Okubo , Haoling Xiong , Atsushi Shimada

Capturing voxel-wise spatial correspondence across distinct modalities is crucial for medical image analysis. However, current registration approaches are not practical enough in terms of registration accuracy and clinical applicability. In…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Tao Guo , Yinuo Wang , Shihao Shu , Weimin Yuan , Diansheng Chen , Zhouping Tang , Cai Meng , Xiangzhi Bai

The demand for lightweight models in image classification tasks under resource-constrained environments necessitates a balance between computational efficiency and robust feature representation. Traditional attention mechanisms, despite…

机器学习 · 计算机科学 2025-04-21 Zhenkai Qin , Feng Zhu , Huan Zeng , Xunyi Nong

Multi-modal recommendation greatly enhances the performance of recommender systems by modeling the auxiliary information from multi-modality contents. Most existing multi-modal recommendation models primarily exploit multimedia information…

信息检索 · 计算机科学 2024-07-09 Xinglong Wu , Anfeng Huang , Hongwei Yang , Hui He , Yu Tai , Weizhe Zhang

Mamba-based architectures have shown to be a promising new direction for deep learning models owing to their competitive performance and sub-quadratic deployment speed. However, current Mamba multi-modal large language models (MLLM) are…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Yifei Xing , Xiangyuan Lan , Ruiping Wang , Dongmei Jiang , Wenjun Huang , Qingfang Zheng , Yaowei Wang

Recent years have seen impressive progress in visual recognition on many benchmarks, however, generalization to the real-world in out-of-distribution setting remains a significant challenge. A state-of-the-art method for robust visual…

计算机视觉与模式识别 · 计算机科学 2021-11-29 Sebastian Cygert , Andrzej Czyzewski

Sufficient cross-task interaction is crucial for success in multi-task dense prediction. However, sufficient interaction often results in high computational complexity, forcing existing methods to face the trade-off between interaction…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Mang Cao , Sanping Zhou , Yizhe Li , Ye Deng , Wenli Huang , Le Wang

Multimodal object detection offers a promising prospect to facilitate robust detection in various visual conditions. However, existing two-stream backbone networks are challenged by complex fusion and substantial parameter increments. This…

计算机视觉与模式识别 · 计算机科学 2024-07-24 Weiying Xie , Yusi Zhang , Tianlin Hui , Jiaqing Zhang , Jie Lei , Yunsong Li

Multimodal Large Language Models (MLLMs) have significantly advanced AI-assisted medical diagnosis, but they often generate factually inconsistent responses that deviate from established medical knowledge. Retrieval-Augmented Generation…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Jinhong Wang , Tajamul Ashraf , Zongyan Han , Jorma Laaksonen , Rao Mohammad Anwer