English
Related papers

Related papers: Are Vision Foundation Models Foundational for Elec…

200 papers

Recently, we have observed that Large Multi-modal Models (LMMs) are revolutionizing the way machines interact with the world, unlocking new possibilities across various multi-modal applications. To adapt LMMs for downstream tasks,…

Computation and Language · Computer Science 2024-11-04 Donghoon Kim , Gusang Lee , Kyuhong Shim , Byonghyo Shim

Vision foundation models have achieved remarkable progress across various image analysis tasks. In the image segmentation task, foundation models like the Segment Anything Model (SAM) enable generalizable zero-shot segmentation through…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Xingxin He , Yifan Hu , Zhaoye Zhou , Mohamed Jarraya , Fang Liu

The use of Environmental Microorganisms (EMs) offers a highly efficient, low cost and harmless remedy to environmental pollution, by monitoring and decomposing of pollutants. This relies on how the EMs are correctly segmented and…

Computer Vision and Pattern Recognition · Computer Science 2022-09-01 Frank Kulwa , Chen Li , Marcin Grzegorzek , Md Mamunur Rahaman , Kimiaki Shirahama , Sergey Kosov

Parameter-efficient fine-tuning (PEFT) has become a popular way to adapt large pre-trained models to new tasks. Most PEFT methods update only a small subset of parameters while freezing the rest, avoiding redundant computation. As they…

Machine Learning · Computer Science 2025-08-25 Sungmin Kang , Jisoo Kim , Salman Avestimehr , Sunwoo Lee

A common strategy for Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViTs) involves adapting the model to downstream tasks by learning a low-rank adaptation matrix. This matrix is decomposed into a product of…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Wei Dong , Yuan Sun , Yiting Yang , Xing Zhang , Zhijun Lin , Qingsen Yan , Haokui Zhang , Peng Wang , Yang Yang , Hengtao Shen

Vision foundation models (FMs) have become the predominant architecture in computer vision, providing highly transferable representations learned from large-scale, multimodal corpora. Nonetheless, they exhibit persistent limitations on…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Fatemeh Ziaeetabar

The Segment Anything Model (SAM) has emerged as a powerful visual foundation model for image segmentation. However, adapting SAM to specific downstream tasks, such as medical and agricultural imaging, remains a significant challenge. To…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Renqi Chen , Haoyang Su , Shixiang Tang

Adapting foundation models for medical image analysis requires finetuning them on a considerable amount of data because of extreme distribution shifts between natural (source) data used for pretraining and medical (target) data. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-01 Mothilal Asokan , Joseph Geo Benjamin , Mohammad Yaqub , Karthik Nandakumar

Medical image segmentation is crucial for clinical diagnosis. The Segmentation Anything Model (SAM) serves as a powerful foundation model for visual segmentation and can be adapted for medical image segmentation. However, medical imaging…

Image and Video Processing · Electrical Eng. & Systems 2024-11-07 Yuxi Liu , Guibo Luo , Yuesheng Zhu

Parameter-efficient fine-tuning (PEFT) reduces the training cost of full-parameter fine-tuning for large language models (LLMs) by training only a small set of task-specific parameters while freezing the pretrained backbone. However,…

Computation and Language · Computer Science 2026-04-22 Xianming Li , Zongxi Li , Tsz-fung Andrew Lee , Jing Li , Haoran Xie , Qing Li

Foundation models have become prominent in computer vision, achieving notable success in various tasks. However, their effectiveness largely depends on pre-training with extensive datasets. Applying foundation models directly to small…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Bowen Zhang , Ying Chen , Long Bai , Yan Zhao , Yuxiang Sun , Yixuan Yuan , Jianhua Zhang , Hongliang Ren

Stereo matching has become a key technique for 3D environment perception in intelligent vehicles. For a considerable time, convolutional neural networks (CNNs) have remained the mainstream choice for feature extraction in this domain.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Chuang-Wei Liu , Qijun Chen , Rui Fan

Recent advancements in machine learning (ML) and deep learning (DL), particularly through the introduction of Foundation Models (FMs), have significantly enhanced surgical scene understanding within minimally invasive surgery (MIS). This…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Ufaq Khan , Umair Nawaz , Adnan Qayyum , Shazad Ashraf , Yutong Xie , Muhammad Haris Khan , Muhammad Bilal , Junaid Qadir

The success of large language models has inspired the computer vision community to explore image segmentation foundation model that is able to zero/few-shot generalize through prompt engineering. Segment-Anything(SAM), among others, is the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Haojie Zhang , Yongyi Su , Xun Xu , Kui Jia

The performance of deep learning models is known to scale with data quantity and diversity. In pathology, as in many other medical imaging domains, the availability of labeled images for a specific task is often limited. Self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Jonas Ammeling , Jonathan Ganz , Emely Rosbach , Ludwig Lausser , Christof A. Bertram , Katharina Breininger , Marc Aubreville

As FMs drive progress toward Artificial General Intelligence (AGI), fine-tuning them under privacy and resource constraints has become increasingly critical particularly when highquality training data resides on distributed edge devices.…

Machine Learning · Computer Science 2025-08-27 Gang Hu , Yinglei Teng , Pengfei Wu , Nan Wang

Accurate 3D mitochondria instance segmentation in electron microscopy (EM) is a challenging problem and serves as a prerequisite to empirically analyze their distributions and morphology. Most existing approaches employ 3D convolutions to…

Image and Video Processing · Electrical Eng. & Systems 2023-03-22 Omkar Thawakar , Rao Muhammad Anwer , Jorma Laaksonen , Orly Reiner , Mubarak Shah , Fahad Shahbaz Khan

Parameter-efficient fine-tuning (PEFT) of pre-trained foundation models is increasingly attracting interest in medical imaging due to its effectiveness and computational efficiency. Among these methods, Low-Rank Adaptation (LoRA) is a…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Ghassen Baklouti , Julio Silva-Rodríguez , Jose Dolz , Houda Bahig , Ismail Ben Ayed

Scene-level neural volumetric reconstruction from monocular videos remains challenging, especially under severe domain shifts. Although recent advances in vision foundation models (VFMs) provide transferable generalized priors learned from…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Yuhang Ming , Tingkang Xi , Xingrui Yang , Lixin Yang , Yong Peng , Cewu Lu , Wanzeng Kong

The automatic diagnosis of Parkinson's disease is in high clinical demand due to its prevalence and the importance of targeted treatment. Current clinical practice often relies on diagnostic biomarkers in QSM and NM-MRI images. However, the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Ding Shaodong , Liu Ziyang , Zhou Yijun , Liu Tao