English
Related papers

Related papers: GeoAlignCLIP: Enhancing Fine-Grained Vision-Langua…

200 papers

Early detection of eye diseases like glaucoma, macular degeneration, and diabetic retinopathy is crucial for preventing vision loss. While artificial intelligence (AI) foundation models hold significant promise for addressing these…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Danli Shi , Weiyi Zhang , Jiancheng Yang , Siyu Huang , Xiaolan Chen , Mayinuer Yusufu , Kai Jin , Shan Lin , Shunming Liu , Qing Zhang , Mingguang He

Mixed modality search -- retrieving information across a heterogeneous corpus composed of images, texts, and multimodal documents -- is an important yet underexplored real-world application. In this work, we investigate how contrastive…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Binxu Li , Yuhui Zhang , Xiaohan Wang , Weixin Liang , Ludwig Schmidt , Serena Yeung-Levy

Integrating ground-level geospatial data with rich geographic context, like OpenStreetMap (OSM), into remote sensing (RS) foundation models (FMs) is essential for advancing geospatial intelligence and supporting a broad spectrum of tasks.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Lubian Bai , Xiuyuan Zhang , Siqi Zhang , Zepeng Zhang , Haoyu Wang , Wei Qin , Shihong Du

Object categories are typically organized into a multi-granularity taxonomic hierarchy. When classifying categories at different hierarchy levels, traditional uni-modal approaches focus primarily on image features, revealing limitations in…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Peng Xia , Xingtong Yu , Ming Hu , Lie Ju , Zhiyong Wang , Peibo Duan , Zongyuan Ge

This paper describes Georeference Contrastive Learning of visual Representation (GeoCLR) for efficient training of deep-learning Convolutional Neural Networks (CNNs). The method leverages georeference information by generating a similar…

Computer Vision and Pattern Recognition · Computer Science 2022-06-28 Takaki Yamada , Adam Prügel-Bennett , Stefan B. Williams , Oscar Pizarro , Blair Thornton

Transfer learning enables the sharing of common knowledge among models for a variety of downstream tasks, but traditional methods suffer in limited training data settings and produce narrow models incapable of effectively generalizing under…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Kevin Vogt-Lowell , Noah Lee , Theodoros Tsiligkaridis , Marc Vaillant

This paper addresses the challenge of Granularity Competition in fine-grained classification tasks, which arises due to the semantic gap between multi-granularity labels. Existing approaches typically develop independent hierarchy-aware…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Zhiguang Lu , Qianqian Xu , Shilong Bao , Zhiyong Yang , Qingming Huang

Improving the safety of vision-language models like CLIP via fine-tuning often comes at a steep price, causing significant drops in their generalization performance. We find this trade-off stems from rigid alignment strategies that force…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Adeel Yousaf , Joseph Fioresi , James Beetham , Amrit Singh Bedi , Mubarak Shah

Fine-grained visual classification aims to recognize objects belonging to many subordinate categories of a supercategory, where appearance alone often fails to distinguish highly similar classes. We propose a unified framework that…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Sumit Mamtani , Yash Thesia

In this paper, we demonstrate that CLIP can also be adapted to downstream tasks where its vision-language alignment is suboptimally learned during pre-training on web-crawled data, all without requiring fine-tuning. We explore the case of…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Sohee Kim , Jisu Kang , Dunam Kim , Seokju Lee

Contrastive Language-Image Pre-training (CLIP) has demonstrated strong generalization across a wide range of visual tasks by leveraging large-scale English-image pairs. However, its extension to low-resource languages remains limited due to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Dahyun Chung , Donghyun Shin , Yujin Sung , Seunggi Moon , Jinwoo Jeon , Byung-Jun Lee

Retail product or packaged grocery goods images need to classified in various computer vision applications like self checkout stores, supply chain automation and retail execution evaluation. Previous works explore ways to finetune deep…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Muktabh Mayank Srivastava

Medical image segmentation is a cornerstone of computer-assisted diagnosis and treatment planning. While recent multimodal vision-language models have shown promise in enhancing semantic understanding through textual descriptions, their…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Saivan Talaei , Fatemeh Daneshfar , Abdulhady Abas Abdullah , Mustaqeem Khan

Vision-Language Models (VLMs), such as CLIP, exhibit strong image-text comprehension abilities, facilitating advances in several downstream tasks such as zero-shot image classification, image-text retrieval, and text-to-image generation.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Le Zhang , Rabiul Awal , Aishwarya Agrawal

Multimodal Large Language Models have achieved impressive performance on a variety of vision-language tasks, yet their fine-grained visual perception and precise spatial reasoning remain limited. In this work, we introduce DiG (Differential…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Zhou Tao , Shida Wang , Yongxiang Hua , Haoyu Cao , Linli Xu

The growing availability of co-located geospatial data spanning aerial imagery, street-level views, elevation models, text, and geographic coordinates offers a unique opportunity for multimodal representation learning. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Guillaume Astruc , Eduard Trulls , Jan Hosang , Loic Landrieu , Paul-Edouard Sarlin

The parameter-efficient adaptation of the image-text pretraining model CLIP for video-text retrieval is a prominent area of research. While CLIP is focused on image-level vision-language matching, video-text retrieval demands comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Leqi Shen , Guoqiang Gong , Tianxiang Hao , Tao He , Yifeng Zhang , Pengzhang Liu , Sicheng Zhao , Jungong Han , Guiguang Ding

Remote sensing imagery presents vast, inherently unstructured spatial data, necessitating sophisticated reasoning to interpret complex user intents and contextual relationships beyond simple recognition tasks. In this paper, we aim to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Liang Yao , Fan Liu , Hongbo Lu , Chuanyi Zhang , Rui Min , Shengxiang Xu , Shimin Di , Pai Peng

Geographic information is essential for modeling tasks in fields ranging from ecology to epidemiology. However, extracting relevant location characteristics for a given task can be challenging, often requiring expensive data fusion or…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Konstantin Klemmer , Esther Rolf , Caleb Robinson , Lester Mackey , Marc Rußwurm

Vision-Language Models (VLMs), such as CLIP, have demonstrated impressive zero-shot transfer capabilities in image-level visual perception. However, these models have shown limited performance in instance-level tasks that demand precise…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Lingfeng Yang , Yueze Wang , Xiang Li , Xinlong Wang , Jian Yang