English
Related papers

Related papers: IDEA: Inverted Text with Cooperative Deformable Ag…

200 papers

Person Re-identification (ReID) has been extensively developed for a decade in order to learn the association of images of the same person across non-overlapping camera views. To overcome significant variations between images across camera…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Wei-Shi Zheng , Junkai Yan , Yi-Xing Peng

Identifying fine-grained book genres is essential for enhancing user experience through efficient discovery, personalized recommendations, and improved reader engagement. At the same time, it provides publishers and marketers with valuable…

Information Retrieval · Computer Science 2025-10-21 Utsav Kumar Nareti , Soumi Chattopadhyay , Prolay Mallick , Suraj Kumar , Chandranath Adak , Ayush Vikas Daga , Adarsh Wase , Arjab Roy

Multiple-in-one image restoration (IR) has made significant progress, aiming to handle all types of single degraded image restoration with a single model. However, in real-world scenarios, images often suffer from combinations of multiple…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Yubin Gu , Yuan Meng , Xiaoshuai Sun , Jiayi Ji , Weijian Ruan , Rongrong Ji

Traditional text-based person ReID assumes that person descriptions from witnesses are complete and provided at once. However, in real-world scenarios, such descriptions are often partial or vague. To address this limitation, we introduce a…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Yiding Lu , Mouxing Yang , Dezhong Peng , Peng Hu , Yijie Lin , Xi Peng

Text-to-image person retrieval aims to identify the target person based on a given textual description query. The primary challenge is to learn the mapping of visual and textual modalities into a common latent space. Prior works have…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Ding Jiang , Mang Ye

Image-text retrieval has developed rapidly in recent years. However, it is still a challenge in remote sensing due to visual-semantic imbalance, which leads to incorrect matching of non-semantic visual and textual features. To solve this…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Qing Ma , Jiancheng Pan , Cong Bai

With the rapid advancement of remote sensing technology, high-resolution multi-modal imagery is now more widely accessible. Conventional Object detection models are trained on a single dataset, often restricted to a specific imaging…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Yuxuan Li , Xiang Li , Yunheng Li , Yicheng Zhang , Yimian Dai , Qibin Hou , Ming-Ming Cheng , Jian Yang

Image-sentence retrieval has attracted extensive research attention in multimedia and computer vision due to its promising application. The key issue lies in jointly learning the visual and textual representation to accurately estimate…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Xuri Ge , Fuhai Chen , Songpei Xu , Fuxiang Tao , Joemon M. Jose

Text-to-image person re-identification (TIReID) retrieves pedestrian images of the same identity based on a query text. However, existing methods for TIReID typically treat it as a one-to-one image-text matching problem, only focusing on…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 Shuanglin Yan , Neng Dong , Jun Liu , Liyan Zhang , Jinhui Tang

The rapid advancement of the automotive industry towards automated and semi-automated vehicles has rendered traditional methods of vehicle interaction, such as touch-based and voice command systems, inadequate for a widening range of…

Human-Computer Interaction · Computer Science 2024-02-08 Amr Gomaa , Guillermo Reyes , Michael Feld , Antonio Krüger

Multimodal emotion recognition in conversation (MERC) requires representations that effectively integrate signals from multiple modalities. These signals include modality-specific cues, information shared across modalities, and interactions…

Machine Learning · Computer Science 2026-01-22 Anh-Tuan Mai , Cam-Van Thi Nguyen , Duc-Trong Le

Multimodal Model Editing (MMED) aims to correct erroneous knowledge in multimodal models. Existing evaluation methods, adapted from textual model editing, overstate success by relying on low-similarity or random inputs, obscure overfitting.…

Machine Learning · Computer Science 2025-11-18 Xiaoqi Han , Ru Li , Ran Yi , Hongye Tan , Zhuomin Liang , Víctor Gutiérrez-Basulto , Jeff Z. Pan

Previous works on multi-label image recognition (MLIR) usually use CNNs as a starting point for research. In this paper, we take pure Vision Transformer (ViT) as the research base and make full use of the advantages of Transformer with…

Computer Vision and Pattern Recognition · Computer Science 2022-04-25 Yunqing Hu , Xuan Jin , Yin Zhang , Haiwen Hong , Jingfeng Zhang , Feihu Yan , Yuan He , Hui Xue

Multi-modal intent detection aims to utilize various modalities to understand the user's intentions, which is essential for the deployment of dialogue systems in real-world scenarios. The two core challenges for multi-modal intent detection…

Computation and Language · Computer Science 2024-01-02 Shijue Huang , Libo Qin , Bingbing Wang , Geng Tu , Ruifeng Xu

Generalizable vehicle re-identification (ReID) seeks to develop models that can adapt to unknown target domains without the need for additional fine-tuning or retraining. Previous works have mainly focused on extracting domain-invariant…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Zhenyu Kuang , Hongyang Zhang , Mang Ye , Bin Yang , Yinhao Liu , Yue Huang , Xinghao Ding , Huafeng Li

The diversity of retinal imaging devices poses a significant challenge: domain shift, which leads to performance degradation when applying the deep learning models trained on one domain to new testing domains. In this paper, we propose a…

Image and Video Processing · Electrical Eng. & Systems 2021-10-07 Peng Liu , Charlie T. Tran , Bin Kong , Ruogu Fang

The ability to jointly learn from multiple modalities, such as text, audio, and visual data, is a defining feature of intelligent systems. While there have been promising advances in designing neural networks to harness multimodal data, the…

Machine Learning · Computer Science 2023-04-25 Zichang Liu , Zhiqiang Tang , Xingjian Shi , Aston Zhang , Mu Li , Anshumali Shrivastava , Andrew Gordon Wilson

Understanding high-resolution (HR) images remains a critical challenge for multimodal large language models (MLLMs). Recent approaches leverage vision-based retrieval-augmented generation (RAG) to retrieve query-relevant crops from HR…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Fan Yang , Xingping Dong , Xin Yu , Wenhan Luo , Wei Liu , Kaihao Zhang

Object detection has been extensively utilized in autonomous systems in recent years, encompassing both 2D and 3D object detection. Recent research in this field has primarily centered around multimodal approaches for addressing this…

Computer Vision and Pattern Recognition · Computer Science 2023-04-20 Wendong Zhang

Text-to-image person re-identification (TIReID) aims to retrieve person images from a large gallery given free-form textual descriptions. TIReID is challenging due to the substantial modality gap between visual appearances and textual…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Giyeol Kim , Chanho Eom
‹ Prev 1 8 9 10 Next ›