中文
相关论文

相关论文: Generalist versus Specialist Vision Foundation Mod…

200 篇论文

The advent of foundation models (FMs) is transforming medical domain. In ophthalmology, RETFound, a retina-specific FM pre-trained sequentially on 1.4 million natural images and 1.6 million retinal images, has demonstrated high adaptability…

The integration of deep learning systems into healthcare has been hindered by the resource-intensive process of data annotation and the inability of these systems to generalize to different data distributions. Foundation models, which are…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Mohammed Baharoon , Waseem Qureshi , Jiahong Ouyang , Yanwu Xu , Abdulrhman Aljouie , Wei Peng

Large vision foundation models have been widely adopted for retinal disease classification without systematic evidence justifying their parameter requirements. In the present work we address two critical questions: First, are large…

图像与视频处理 · 电气工程与系统科学 2025-12-01 David Isztl , Tahm Spitznagel , Gabor Mark Somfai , Rui Santos

The advent of large-scale vision foundation models, pre-trained on diverse natural images, has marked a paradigm shift in computer vision. However, how the frontier vision foundation models' efficacies transfer to specialised domains such…

Vision foundation models like DINOv2 demonstrate remarkable potential in medical imaging despite their origin in natural image domains. However, their design inherently works best for uni-modal image analysis, limiting their effectiveness…

图像与视频处理 · 电气工程与系统科学 2025-09-09 Daniel Scholz , Ayhan Can Erdur , Viktoria Ehm , Anke Meyer-Baese , Jan C. Peeken , Daniel Rueckert , Benedikt Wiestler

Foundation vision encoders such as CLIP and DINOv2, trained on web-scale data, exhibit strong transfer performance across tasks and datasets. However, medical imaging foundation models remain constrained by smaller datasets, limiting our…

Background: RETFound, a self-supervised, retina-specific foundation model (FM), showed potential in downstream applications. However, its comparative performance with traditional deep learning (DL) models remains incompletely understood.…

Integrating deep learning into medical imaging is poised to greatly advance diagnostic methods but it faces challenges with generalizability. Foundation models, based on self-supervised learning, address these issues and improve data…

The deep learning field is converging towards the use of general foundation models that can be easily adapted for diverse tasks. While this paradigm shift has become common practice within the field of natural language processing, progress…

计算机视觉与模式识别 · 计算机科学 2023-11-15 Joana Palés Huix , Adithya Raju Ganeshan , Johan Fredin Haslum , Magnus Söderberg , Christos Matsoukas , Kevin Smith

Using massive datasets, foundation models are large-scale, pre-trained models that perform a wide range of tasks. These models have shown consistently improved results with the introduction of new methods. It is crucial to analyze how these…

图像与视频处理 · 电气工程与系统科学 2025-05-27 Mobina Mansoori , Sajjad Shahabodini , Farnoush Bayatmakou , Jamshid Abouei , Konstantinos N. Plataniotis , Arash Mohammadi

With access to large-scale, unlabeled medical datasets, researchers are confronted with two questions: Should they attempt to pretrain a custom foundation model on this medical data, or use transfer-learning from an existing generalist…

Vision foundation models have attracted significant attention for their ability to leverage large-scale unlabeled visual data. This advantage is particularly important in remote sensing, where data acquisition is costly and annotation often…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Hyobin Park , Minseok Seo , Dong-Geol Choi

Foundation models are predominantly trained in an unsupervised or self-supervised manner on highly diverse and large-scale datasets, making them broadly applicable to various downstream tasks. In this work, we investigate for the first time…

计算机视觉与模式识别 · 计算机科学 2025-02-10 Tahar Chettaoui , Naser Damer , Fadi Boutros

Foundation models have become prominent in computer vision, achieving notable success in various tasks. However, their effectiveness largely depends on pre-training with extensive datasets. Applying foundation models directly to small…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Bowen Zhang , Ying Chen , Long Bai , Yan Zhao , Yuxiang Sun , Yixuan Yuan , Jianhua Zhang , Hongliang Ren

Artificial intelligence (AI) is vital in ophthalmology, tackling tasks like diagnosis, classification, and visual question answering (VQA). However, existing AI models in this domain often require extensive annotation and are task-specific,…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Danli Shi , Weiyi Zhang , Xiaolan Chen , Yexin Liu , Jiancheng Yang , Siyu Huang , Yih Chung Tham , Yingfeng Zheng , Mingguang He

Learning-based monocular visual odometry (VO) poses robustness, generalization, and efficiency challenges in robotics. Recent advances in visual foundation models, such as DINOv2, have improved robustness and generalization in various…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Maulana Bisyir Azhari , David Hyunchul Shim

Large-scale vision foundation models such as DINOv2 boast impressive performances by leveraging massive architectures and training datasets. But numerous scenarios require practitioners to reproduce those pre-training solutions, such as on…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Jiaqi Zhang , Juntuo Wang , Zhixin Sun , John Zou , Randall Balestriero

Medical image analysis frequently encounters data scarcity challenges. Transfer learning has been effective in addressing this issue while conserving computational resources. The recent advent of foundational models like the DINOv2, which…

图像与视频处理 · 电气工程与系统科学 2024-02-14 Yuning Huang , Jingchen Zou , Lanxi Meng , Xin Yue , Qing Zhao , Jianqiang Li , Changwei Song , Gabriel Jimenez , Shaowu Li , Guanghui Fu

Accurate segmentation of organs and tumors in CT and MRI scans is essential for diagnosis, treatment planning, and disease monitoring. While deep learning has advanced automated segmentation, most models remain task-specific, lacking…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Yuheng Li , Yizhou Wu , Yuxiang Lai , Mingzhe Hu , Xiaofeng Yang

Face Anti-Spoofing (FAS) remains challenging due to the requirement for robust domain generalization across unseen environments. While recent trends leverage Vision-Language Models (VLMs) for semantic supervision, these multimodal…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Mika Feng , Pierre Gallin-Martel , Koichi Ito , Takafumi Aoki
‹ 上一页 1 2 3 10 下一页 ›