English
Related papers

Related papers: Cross-Modal Pre-Aligned Method with Global and Loc…

200 papers

Vision-language alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize mutual information, primarily aligning pairwise samples…

Machine Learning · Computer Science 2026-02-25 Wenzhe Yin , Zehao Xiao , Pan Zhou , Shujian Yu , Jiayi Shen , Jan-Jakob Sonke , Efstratios Gavves

Multi-modal object Re-IDentification (ReID) is devoted to retrieving specific objects through the exploitation of complementary multi-modal image information. Existing methods mainly concentrate on the fusion of multi-modal features, yet…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Yangyang Liu , Yuhao Wang , Pingping Zhang

Multimodal learning seeks to integrate information across diverse sensory sources, yet current approaches struggle to balance cross-modal generalizability with modality-specific structure. Continuous (implicit) methods preserve fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Souptik Sen , Raneen Younis , Zahra Ahmadi

We introduce a method for manifold alignment of different modalities (or domains) of remote sensing images. The problem is recurrent when a set of multitemporal, multisource, multisensor and multiangular images is available. In these…

Computer Vision and Pattern Recognition · Computer Science 2021-04-19 Devis Tuia , Michele Volpi , Maxime Trolliet , Gustau Camps-Valls

Remote sensing (RS) images are usually stored in compressed format to reduce the storage size of the archives. Thus, existing content-based image retrieval (CBIR) systems in RS require decoding images before applying CBIR (which is…

Computer Vision and Pattern Recognition · Computer Science 2023-01-24 Gencer Sumbul , Jun Xiang , Nimisha Thekke Madam , Begüm Demir

Image-to-point cloud cross-modal Visual Place Recognition (VPR) is a challenging task where the query is an RGB image, and the database samples are LiDAR point clouds. Compared to single-modal VPR, this approach benefits from the widespread…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Jianyi Peng , Fan Lu , Bin Li , Yuan Huang , Sanqing Qu , Guang Chen

Multimodal object detection has attracted significant attention in both academia and industry for its enhanced robustness. Although numerous studies have focused on improving modality fusion strategies, most neglect fusion degradation, and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 YiKang Shao , Tao Shi

Image-sentence retrieval has attracted extensive research attention in multimedia and computer vision due to its promising application. The key issue lies in jointly learning the visual and textual representation to accurately estimate…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Xuri Ge , Fuhai Chen , Songpei Xu , Fuxiang Tao , Joemon M. Jose

Image-point class incremental learning helps the 3D-points-vision robots continually learn category knowledge from 2D images, improving their perceptual capability in dynamic environments. However, some incremental learning methods address…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Chao Qi , Jianqin Yin , Ren Zhang

In real-world applications of human pose estimation, low-resolution input images are frequently encountered when the performance of the image acquisition equipment is limited or the shooting distance is too far. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Zejun Gu , Zhong-Qiu Zhao , Hao Shen , Zhao Zhang

We focus on domain and class generalization problems in analyzing optical remote sensing images, using the large-scale pre-trained vision-language model (VLM), CLIP. While contrastively trained VLMs show impressive zero-shot generalization…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Avigyan Bhattacharya , Mainak Singha , Ankit Jha , Biplab Banerjee

Cross-Modal Retrieval (CMR) is an important research topic across multimodal computing and information retrieval, which takes one type of data as the query to retrieve relevant data of another type. It has been widely used in many…

Computer Vision and Pattern Recognition · Computer Science 2022-04-19 Zhixiong Zeng , Wenji Mao

Pre-training plays a vital role in various vision tasks, such as object recognition and detection. Commonly used pre-training methods, which typically rely on randomized approaches like uniform or Gaussian distributions to initialize model…

Computer Vision and Pattern Recognition · Computer Science 2024-11-15 Chen-Long Duan , Yong Li , Xiu-Shen Wei , Lin Zhao

Vision-language alignment in multi-modal large language models (MLLMs) relies on supervised fine-tuning (SFT) or reinforcement learning (RL). To align multi-modal large language models (MLLMs) in the post-training stage, supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Xin Jin , Siyuan Li , Siyong Jian , Kai Yu , Huan Wang

Text-based person re-identification(Re-id) is an important task in video surveillance, which consists of retrieving the corresponding person's image given a textual description from a large gallery of images. It is difficult to directly…

Computer Vision and Pattern Recognition · Computer Science 2020-09-22 Tinghuai Ma , Mingming Yang , Huan Rong , Yurong Qian , Yurong Qian , Yuan Tian , NajlaAl-Nabhan

Image fusion integrates complementary information from different modalities to generate high-quality fused images, thereby enhancing downstream tasks such as object detection and semantic segmentation. Unlike task-specific techniques that…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Yingying Wang , Rongjin Zhuang , Hui Zheng , Xuanhua He , Ke Cao , Xiaotong Tu , Xinghao Ding

Multimodal fusion has made great progress in the field of remote sensing image classification due to its ability to exploit the complementary spatial-spectral information. Deep learning methods such as CNN and Transformer have been widely…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Qingyu Wang , Xue Jiang , Guozheng Xu

This paper presents an innovative framework for remote sensing image analysis by fusing deep learning techniques, specifically Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks, with Geographic Information…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Sajjad Afroosheh , Mohammadreza Askari

Most existing methods in vision-language retrieval match two modalities by either comparing their global feature vectors which misses sufficient information and lacks interpretability, detecting objects in images or videos and aligning the…

Computer Vision and Pattern Recognition · Computer Science 2022-10-04 Xiaohan Zou , Changqiao Wu , Lele Cheng , Zhongyuan Wang

Text-to-image person retrieval aims to identify the target person based on a given textual description query. The primary challenge is to learn the mapping of visual and textual modalities into a common latent space. Prior works have…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Ding Jiang , Mang Ye