English
Related papers

Related papers: MOS: Mitigating Optical-SAR Modality Gap for Cross…

200 papers

This paper introduces reconstructive visual instruction tuning (ROSS), a family of Large Multimodal Models (LMMs) that exploit vision-centric supervision signals. In contrast to conventional visual instruction tuning approaches that…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Haochen Wang , Anlin Zheng , Yucheng Zhao , Tiancai Wang , Zheng Ge , Xiangyu Zhang , Zhaoxiang Zhang

Referring image segmentation (RIS) is a fundamental vision-language task that intends to segment a desired object from an image based on a given natural language expression. Due to the essentially distinct data properties between image and…

Computer Vision and Pattern Recognition · Computer Science 2024-02-15 Wenxuan Wang , Jing Liu , Xingjian He , Yisi Zhang , Chen Chen , Jiachen Shen , Yan Zhang , Jiangyun Li

Cross-modality person re-identification (cm-ReID) is a challenging but key technology for intelligent video analysis. Existing works mainly focus on learning common representation by embedding different modalities into a same feature space.…

Computer Vision and Pattern Recognition · Computer Science 2020-03-13 Yan Lu , Yue Wu , Bin Liu , Tianzhu Zhang , Baopu Li , Qi Chu , Nenghai Yu

Object Re-Identification (ReID) is pivotal in computer vision, witnessing an escalating demand for adept multimodal representation learning. Current models, although promising, reveal scalability limitations with increasing modalities as…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Haoli Yin , Jiayao Li , Eva Schiller , Luke McDermott , Daniel Cummings

Unsupervised visible-infrared person re-identification (USL-VI-ReID) endeavors to retrieve pedestrian images of the same identity from different modalities without annotations. While prior work focuses on establishing cross-modality…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Lingfeng He , De Cheng , Nannan Wang , Xinbo Gao

Referring remote sensing image segmentation (RRSIS) is a novel visual task in remote sensing images segmentation, which aims to segment objects based on a given text description, with great significance in practical application. Previous…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Leideng Shi , Juan Zhang

Referring image segmentation (RIS) aims to segment a particular region based on a language expression prompt. Existing methods incorporate linguistic features into visual features and obtain multi-modal features for mask decoding. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Mengxi Zhang , Yiming Liu , Xiangjun Yin , Huanjing Yue , Jingyu Yang

Visible-Infrared person re-identification (VI-ReID) is an important and challenging task in intelligent video surveillance. Existing methods mainly focus on learning a shared feature space to reduce the modality discrepancy between visible…

Computer Vision and Pattern Recognition · Computer Science 2023-08-30 Haichao Shi , Mandi Luo , Xiao-Yu Zhang , Ran He

Multimodal learning leverages complementary information derived from different modalities, thereby enhancing performance in medical image segmentation. However, prevailing multimodal learning methods heavily rely on extensive well-annotated…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Xiaogen Zhou , Yiyou Sun , Min Deng , Winnie Chiu Wing Chu , Qi Dou

RGB-Infrared person re-identification (RGB-IR Re-ID) aims to match persons from heterogeneous images captured by visible and thermal cameras, which is of great significance in the surveillance system under poor light conditions. Facing…

Computer Vision and Pattern Recognition · Computer Science 2020-12-15 Can Zhang , Hong Liu , Wei Guo , Mang Ye

Remote Sensing Image-Text Retrieval (RSITR) is pivotal for knowledge services and data mining in the remote sensing (RS) domain. Considering the multi-scale representations in image content and text vocabulary can enable the models to learn…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Rui Yang , Shuang Wang , Yingping Han , Yuanheng Li , Dong Zhao , Dou Quan , Yanhe Guo , Licheng Jiao

Multimodal Magnetic Resonance Imaging (MRI) provides essential complementary information for analyzing brain tumor subregions. While methods using four common MRI modalities for automatic segmentation have shown success, they often face…

Image and Video Processing · Electrical Eng. & Systems 2024-11-14 Runze Cheng , Zhongao Sun , Ye Zhang , Chun Li

Both image registration and label fusion in the multi-atlas segmentation (MAS) rely on the intensity similarity between target and atlas images. However, such similarity can be problematic when target and atlas images are acquired using…

Computer Vision and Pattern Recognition · Computer Science 2020-08-21 Wangbin Ding , Lei Li , Xiahai Zhuang , Liqin Huang

Due to its potential wide applications in video surveillance and other computer vision tasks like tracking, person re-identification (ReID) has become popular and been widely investigated. However, conventional person re-identification can…

Computer Vision and Pattern Recognition · Computer Science 2020-03-03 Xing Fan , Hao Luo , Chi Zhang , Wei Jiang

Multi-modal fusion is imperative to the implementation of reliable object detection and tracking in complex environments. Exploiting the synergy of heterogeneous modal information endows perception systems the ability to achieve more…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Kun Shi , Shibo He , Zhenyu Shi , Anjun Chen , Zehui Xiong , Jiming Chen , Jun Luo

Cross-modal retrieval is the task of retrieving samples of a given modality by using queries of a different one. Due to the wide range of practical applications, the problem has been mainly focused on the vision and language case, e.g. text…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Jorge Sánchez , Rodrigo Laguna

Multimodal industrial surface defect detection (MISDD) aims to identify and locate defect in industrial products by fusing RGB and 3D modalities. This article focuses on modality-missing problems caused by uncertain sensors availability in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Shuai Jiang , Yunfeng Ma , Jingyu Zhou , Yuan Bian , Yaonan Wang , Min Liu

The development of cross-modal retrieval systems that can search and retrieve semantically relevant data across different modalities based on a query in any modality has attracted great attention in remote sensing (RS). In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2022-04-20 Georgii Mikriukov , Mahdyar Ravanbakhsh , Begüm Demir

The Visible-Infrared Person Re-identification (VI ReID) aims to match visible and infrared images of the same pedestrians across non-overlapped camera views. These two input modalities contain both invariant information, such as shape, and…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Ruiqi Wu , Bingliang Jiao , Wenxuan Wang , Meng Liu , Peng Wang

Remote sensing image interpretation plays a critical role in environmental monitoring, urban planning, and disaster assessment. However, acquiring high-quality labeled data is often costly and time-consuming. To address this challenge, we…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Tong Wang , Guanzhou Chen , Xiaodong Zhang , Chenxi Liu , Jiaqi Wang , Xiaoliang Tan , Wenchao Guo , Qingyuan Yang , Kaiqi Zhang
‹ Prev 1 3 4 5 6 7 10 Next ›