English
Related papers

Related papers: FRoundation: Are Foundation Models Ready for Face …

200 papers

The recent development of foundation models for time series data has generated considerable interest in using such models across a variety of applications. Although foundation models achieve state-of-the-art predictive performance, their…

Machine Learning · Computer Science 2026-05-29 Coen Adler , Yuxin Chang , Felix Draxler , Samar Abdi , Padhraic Smyth

In the present work, we show that the performance of formula-driven supervised learning (FDSL) can match or even exceed that of ImageNet-21k and can approach that of the JFT-300M dataset without the use of real images, human supervision, or…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Hirokatsu Kataoka , Sora Takashima , Ryo Hayamizu , Ryosuke Yamada , Kodai Nakashima , Xinyu Zhang , Edgar Josafat Martinez-Noriega , Nakamasa Inoue , Rio Yokota

Due to the increase in computational resources and accessibility of data, an increase in large, deep learning models trained on copious amounts of multi-modal data using self-supervised or semi-supervised learning have emerged. These…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Madeline Chantry Schiappa , Shehreen Azad , Sachidanand VS , Yunhao Ge , Ondrej Miksik , Yogesh S. Rawat , Vibhav Vineet

Face recognition (FR) models are vulnerable to performance variations across demographic groups. The causes for these performance differences are unclear due to the highly complex deep learning-based structure of face recognition models.…

Computer Vision and Pattern Recognition · Computer Science 2025-01-29 Marco Huber , Fadi Boutros , Naser Damer

Clinical deployment of automated brain MRI analysis faces a fundamental challenge: clinical data is heterogeneous and noisy, and high-quality labels are prohibitively costly to obtain. Self-supervised learning (SSL) can address this by…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Asbjørn Munk , Stefano Cerri , Vardan Nersesjan , Christian Hedeager Krag , Jakob Ambsdorf , Pablo Rocamora García , Julia Machnio , Peirong Liu , Suhyun Ahn , Nasrin Akbari , Yasmina Al Khalil , Kimberly Amador , Sina Amirrajab , Tal Arbel , Meritxell Bach Cuadra , Ujjwal Baid , Bhakti Baheti , Jaume Banus , Kamil Barbierik , Christoph Brune , Yansong Bu , Baptiste Callard , Yuhan Chen , Cornelius Crijnen , Corentin Dancette , Peter Drotar , Prasad Dutande , Nils D. Forkert , Saurabh Garg , Jakub Gazda , Matej Gazda , Benoît Gérin , Partha Ghosh , Weikang Gong , Pedro M. Gordaliza , Sam Hashemi , Tobias Heimann , Fucang Jia , Jiexin Jiang , Emily Kaczmarek , Chris Kang , Seung Kwan Kang , Mohammad Khazaei , Julien Khlaut , Petros Koutsouvelis , Jae Sung Lee , Yuchong Li , Mengye Lyu , Mingchen Ma , Anant Madabhushi , Klaus H. Maier-Hein , Pierre Manceron , Andrés Martínez Mora , Moona Mazher , Felix Meister , Nataliia Molchanova , Steven A. Niederer , Leonard Nürnberg , Jinah Park , Abdul Qayyum , Jonas Richiardi , Antoine Saporta , Branislav Setlak , Ning Shen , Justin Szeto , Constantin Ulrich , Puru Vaish , Vibujithan Vigneshwaran , Leroy Volmer , Zihao Wang , Siqi Wei , Anthony Winder , Jelmer M. Wolterink , Maxence Wynen , Chang Yang , Si Young Yie , Mostafa Mehdipour Ghazi , Akshay Pai , Espen Jimenez Solem , Sebastian Nørgaard Llambias , Mikael Boesen , Michael Eriksen Benros , Juan Eugenio Iglesias , Mads Nielsen

Birds Eye View perception models require extensive data to perform and generalize effectively. While traditional datasets often provide abundant driving scenes from diverse locations, this is not always the case. It is crucial to maximize…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Seamie Hayes , Ganesh Sistu , Ciarán Eising

Foundation models have revolutionized artificial intelligence by providing robust, versatile architectures pre-trained on large-scale datasets. However, adapting these massive models to specific downstream tasks requires fine-tuning, which…

Machine Learning · Computer Science 2025-05-01 Jieming Bian , Yuanzhe Peng , Lei Wang , Yin Huang , Jie Xu

Over the past years, deep learning capabilities and the availability of large-scale training datasets advanced rapidly, leading to breakthroughs in face recognition accuracy. However, these technologies are foreseen to face a major…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Fadi Boutros , Vitomir Struc , Julian Fierrez , Naser Damer

Frozen pretrained models have become a viable alternative to the pretraining-then-finetuning paradigm for transfer learning. However, with frozen models there are relatively few parameters available for adapting to downstream tasks, which…

Computer Vision and Pattern Recognition · Computer Science 2022-11-04 Yutong Lin , Ze Liu , Zheng Zhang , Han Hu , Nanning Zheng , Stephen Lin , Yue Cao

Foundation Models (FMs), such as LLaMA, BERT, GPT, ViT, and CLIP, have demonstrated remarkable success in a wide range of applications, driven by their ability to leverage vast amounts of data for pre-training. However, optimizing FMs often…

Machine Learning · Computer Science 2024-03-21 Sixing Yu , J. Pablo Muñoz , Ali Jannesari

Foundation vision encoders such as CLIP and DINOv2, trained on web-scale data, exhibit strong transfer performance across tasks and datasets. However, medical imaging foundation models remain constrained by smaller datasets, limiting our…

Adapting pre-trained models has become an effective strategy in artificial intelligence, offering a scalable and efficient alternative to training models from scratch. In the context of remote sensing (RS), where visual grounding(VG)…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Hasan Moughnieh , Mohamad Chalhoub , Hasan Nasrallah , Cristiano Nattero , Paolo Campanella , Giovanni Nico , Ali J. Ghandour

Existing foundation models (FMs) in the medical domain often require extensive fine-tuning or rely on training resource-intensive decoders, while many existing encoders are pretrained with objectives biased toward specific tasks. This…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Tim Veenboer , George Yiasemis , Eric Marcus , Vivien Van Veldhuizen , Cees G. M. Snoek , Jonas Teuwen , Kevin B. W. Groot Lipman

Pretrained foundation models learn embeddings that can be used for a wide range of downstream tasks. These embeddings optimise general performance, and if insufficiently accurate at a specific task the model can be fine-tuned to improve…

Machine Learning · Computer Science 2025-02-20 Matthew P. Wilson , Edward O. Pyzer-Knapp , Nicolas Galichet , Luke Dicks

Despite the significant progress made by all-in-one models in universal image restoration, existing methods suffer from a generalization bottleneck in real-world scenarios, as they are mostly trained on small-scale synthetic datasets with…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Hao Li , Xiang Chen , Jiangxin Dong , Jinhui Tang , Jinshan Pan

Purpose: Depth estimation in robotic surgery is vital in 3D reconstruction, surgical navigation and augmented reality visualization. Although the foundation model exhibits outstanding performance in many vision tasks, including depth…

Computer Vision and Pattern Recognition · Computer Science 2024-01-15 Beilei Cui , Mobarakol Islam , Long Bai , Hongliang Ren

Foundation models encode rich representations that can be adapted to downstream tasks by fine-tuning. However, fine-tuning a model on one data distribution often degrades performance under distribution shifts. Current approaches to robust…

Machine Learning · Computer Science 2024-03-15 Caroline Choi , Yoonho Lee , Annie Chen , Allan Zhou , Aditi Raghunathan , Chelsea Finn

The small amount of training data for many state-of-the-art deep learning-based Face Recognition (FR) systems causes a marked deterioration in their performance. Although a considerable amount of research has addressed this issue by…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Soroush Hashemifar , Abdolreza Marefat , Javad Hassannataj Joloudari , Hamid Hassanpour

Data scaling has revolutionized research fields like natural language processing, computer vision, and robotics control, providing foundation models with remarkable multi-task and generalization capabilities. In this paper, we investigate…

Systems and Control · Electrical Eng. & Systems 2025-03-27 Shaohuai Liu , Lin Dong , Chao Tian , Le Xie

Foundation models have become prominent in computer vision, achieving notable success in various tasks. However, their effectiveness largely depends on pre-training with extensive datasets. Applying foundation models directly to small…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Bowen Zhang , Ying Chen , Long Bai , Yan Zhao , Yuxiang Sun , Yixuan Yuan , Jianhua Zhang , Hongliang Ren