English
Related papers

Related papers: Learning to Adapt Foundation Model DINOv2 for Caps…

200 papers

The success of large pre-trained object detectors hinges on their adaptability to diverse downstream tasks. While fine-tuning is the standard adaptation method, specializing these models for challenging fine-grained domains necessitates…

Computer Vision and Pattern Recognition · Computer Science 2025-05-05 Vishal Gandhi , Sagar Gandhi

The Parameter-Efficient Fine-Tuning (PEFT) methods have been extensively researched for large language models in downstream tasks. Among all the existing approaches, the Low-Rank Adaptation (LoRA) has gained popularity for its streamlined…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Tangyu Jiang , Haodi Wang , Chun Yuan

Reliable plant species and damage segmentation for herbicide field research trials requires models that can withstand substantial real-world variation across seasons, geographies, devices, and sensing modalities. Most deep learning…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Artzai Picon , Itziar Eguskiza , Daniel Mugica , Javier Romero , Carlos Javier Jimenez , Eric White , Gabriel Do-Lago-Junqueira , Christian Klukas , Ramon Navarra-Mestre

Effectively leveraging private datasets remains a significant challenge in developing foundation models. Federated Learning (FL) has recently emerged as a collaborative framework that enables multiple users to fine-tune these models while…

Machine Learning · Computer Science 2025-10-27 Yiyuan Yang , Guodong Long , Qinghua Lu , Liming Zhu , Jing Jiang , Chengqi Zhang

Foundation models for vision are predominantly trained on RGB data, while many safety-critical applications rely on non-visible modalities such as infrared (IR) and synthetic aperture radar (SAR). We study whether a single flow-matching…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Maxim Clouser , Kia Khezeli , John Kalantari

Recent advances in foundation models have brought promising results in computer vision, including medical image segmentation. Fine-tuning foundation models on specific low-resource medical tasks has become a standard practice. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Jingyun Yang , Guoqing Zhang , Jingge Wang , Yang Li

Recently, many foundation models for medical image analysis such as MedSAM, SwinUNETR have been released and proven to be useful in multiple tasks. However, considering the inherent heterogeneity and inhomogeneity of real-world medical…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Shangde Gao , Yichao Fu , Ke Liu , Hongxia Xu , Jian Wu

Code embeddings are essential for semantic code search; however, current approaches often struggle to capture the precise syntactic and contextual nuances inherent in code. Open-source models such as CodeBERT and UniXcoder exhibit…

Machine Learning · Computer Science 2025-06-03 Saumya Chaturvedi , Aman Chadha , Laurent Bindschaedler

This paper investigates the application of Low-Rank Adaptation (LoRA) to small models for cross-domain few-shot object detection in aerial images. Originally designed for large-scale models, LoRA helps mitigate overfitting, making it a…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Hicham Talaoubrid , Anissa Mokraoui , Ismail Ben Ayed , Axel Prouvost , Sonimith Hang , Monit Korn , Rémi Harvey

Foundation models are predominantly trained in an unsupervised or self-supervised manner on highly diverse and large-scale datasets, making them broadly applicable to various downstream tasks. In this work, we investigate for the first time…

Computer Vision and Pattern Recognition · Computer Science 2025-02-10 Tahar Chettaoui , Naser Damer , Fadi Boutros

Although visual foundation models like DINOv2 provide state-of-the-art performance as feature extractors, their complex, high-dimensional representations create substantial hurdles for interpretability. This work proposes DINO-QPM, which…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Robert Zimmermann , Thomas Norrenbrock , Bodo Rosenhahn

Accurate and privacy-preserving diagnosis of ophthalmic diseases remains a critical challenge in medical imaging, particularly given the limitations of existing deep learning models in handling data imbalance, data privacy concerns, spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Md. Naimur Asif Borno , Md Sakib Hossain Shovon , MD Hanif Sikder , Iffat Firozy Rimi , Tahani Jaser Alahmadi , Mohammad Ali Moni

Processing visual data often involves small adjustments or sequences of changes, e.g., image filtering, surface smoothing, and animation. While established graphics techniques like normal mapping and video compression exploit redundancy to…

Graphics · Computer Science 2025-10-20 Anh Truong , Ahmed H. Mahmoud , Mina Konaković Luković , Justin Solomon

Vision Language Models (VLMs) integrate visual and text modalities to enable multimodal understanding and generation. These models typically combine a Vision Transformer (ViT) as an image encoder and a Large Language Model (LLM) for text…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Krishna Teja Chitty-Venkata , Murali Emani , Venkatram Vishwanath

In this paper, we introduce Symmetric Low-Rank Adapters, an optimized variant of LoRA with even fewer weights. This method utilizes Low-Rank Symmetric Weight Matrices to learn downstream tasks more efficiently. Traditional LoRA accumulates…

Machine Learning · Computer Science 2025-04-17 Tales Panoutsos , Rodrygo L. T. Santos , Flavio Figueiredo

Large Multimodal Models (LMMs) have shown significant progress in various complex vision tasks with the solid linguistic and reasoning capacity inherited from large language models (LMMs). Low-rank adaptation (LoRA) offers a promising…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Liang Mi , Weijun Wang , Wenming Tu , Qingfeng He , Rui Kong , Xinyu Fang , Yazhu Dong , Yikang Zhang , Yunchun Li , Meng Li , Haipeng Dai , Guihai Chen , Yunxin Liu

As large language models (LLMs) continue to scale in size, the computational overhead has become a major bottleneck for task-specific fine-tuning. While low-rank adaptation (LoRA) effectively curtails this cost by confining the weight…

Machine Learning · Computer Science 2026-05-15 Yilang Zhang , Xiaodong Yang , Yiwei Cai , Georgios B. Giannakis

Endoscopy is a widely used imaging modality to diagnose and treat diseases in hollow organs as for example the gastrointestinal tract, the kidney and the liver. However, due to varied modalities and use of different imaging protocols at…

Computer Vision and Pattern Recognition · Computer Science 2020-03-30 Sharib Ali , Binod Bhattarai , Tae-Kyun Kim , Jens Rittscher

In this paper we generalize and extend an idea of low-rank adaptation (LoRA) of large language models (LLMs) based on Transformer architecture. Widely used LoRA-like methods of fine-tuning LLMs are based on matrix factorization of gradient…

Computation and Language · Computer Science 2024-02-06 Daniel Bershatsky , Daria Cherniuk , Talgat Daulbaev , Aleksandr Mikhalev , Ivan Oseledets

Foundation models have exhibited remarkable success in various applications, such as disease diagnosis and text report generation. To date, a foundation model for endoscopic video analysis is still lacking. In this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Zhao Wang , Chang Liu , Shaoting Zhang , Qi Dou