English
Related papers

Related papers: Benchmarking Vision Foundation Models for Domain-G…

200 papers

Vision Foundation Models (VFMs) are large-scale, pre-trained models that serve as general-purpose backbones for various computer vision tasks. As VFMs' popularity grows, there is an increasing interest in understanding their effectiveness…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Volodymyr Havrylov , Haiwen Huang , Dan Zhang , Andreas Geiger

Self-supervised learning holds the promise of eliminating the need for manual data annotation, enabling models to scale effortlessly to massive datasets and larger architectures. By not being tailored to specific tasks or domains, this…

We propose ADIOS, a masked image model (MIM) framework for self-supervised learning, which simultaneously learns a masking function and an image encoder using an adversarial objective. The image encoder is trained to minimise the distance…

Computer Vision and Pattern Recognition · Computer Science 2022-07-07 Yuge Shi , N. Siddharth , Philip H. S. Torr , Adam R. Kosiorek

Recent vision foundation models (VFMs) have demonstrated proficiency in various tasks but require supervised fine-tuning to perform the task of semantic segmentation effectively. Benchmarking their performance is essential for selecting…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Tommie Kerssies , Daan de Geus , Gijs Dubbelman

In this paper, we describe our submissions to the ZeroSpeech 2021 Challenge and SUPERB benchmark. Our submissions are based on the recently proposed FaST-VGS model, which is a Transformer-based model that learns to associate raw speech…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-03 Puyuan Peng , David Harwath

Language has been useful in extending the vision encoder to data from diverse distributions without empirical discovery in training domains. However, as the image description is mostly at coarse-grained level and ignores visual details, the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Jiawei Ma , Yulei Niu , Shiyuan Huang , Guangxing Han , Shih-Fu Chang

Source-Free Domain Adaptation (SFDA) aims to adapt a source model for a target domain, with only access to unlabeled target training data and the source model pre-trained on a supervised source domain. Relying on pseudo labeling and/or…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Song Tang , Wenxin Su , Mao Ye , Xiatian Zhu

Vision Transformers (ViTs) have revolutionized large-scale visual modeling, yet remain underexplored in face recognition (FR) where CNNs still dominate. We identify a critical bottleneck: CNN-inspired training paradigms fail to unlock ViT's…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Jinghan You , Shanglin Li , Yuanrui Sun , Jiangchuan Wei , Mingyu Guo , Chao Feng , Jiao Ran

Face anti-spoofing (FAS) secures face recognition from presentation attacks (PAs). Existing FAS methods usually supervise PA detectors with handcrafted binary or pixel-wise labels. However, handcrafted labels may are not the most adequate…

Computer Vision and Pattern Recognition · Computer Science 2021-11-15 Yunxiao Qin , Zitong Yu , Longbin Yan , Zezheng Wang , Chenxu Zhao , Zhen Lei

Large-scale contrastive pre-training produces powerful Vision-and-Language Models (VLMs) capable of generating representations (embeddings) effective for a wide variety of visual and multimodal tasks. However, these pretrained embeddings…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Nikolaos-Antonios Ypsilantis , Kaifeng Chen , André Araujo , Ondřej Chum

The challenge of Domain Generalization (DG) in Face Anti-Spoofing (FAS) is the significant interference of domain-specific signals on subtle spoofing clues. Recently, some CLIP-based algorithms have been developed to alleviate this…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Jiabao Guo , Ajian Liu , Yunfeng Diao , Jin Zhang , Hui Ma , Bo Zhao , Richang Hong , Meng Wang

Unsupervised semantic segmentation aims to categorize each pixel in an image into a corresponding class without the use of annotated data. It is a widely researched area as obtaining labeled datasets is expensive. While previous works in…

Computer Vision and Pattern Recognition · Computer Science 2024-01-01 Yau Shing Jonathan Cheung , Xi Chen , Lihe Yang , Hengshuang Zhao

Purpose: This study provides the first comprehensive evaluation of foundation models in fetal ultrasound (US) imaging under low inter-class variability conditions. While recent vision foundation models such as DINOv3 have shown remarkable…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Edoardo Conti , Riccardo Rosati , Lorenzo Federici , Adriano Mancini , Maria Chiara Fiorentin

Face recognition technology has dramatically transformed the landscape of security, surveillance, and authentication systems, offering a user-friendly and non-invasive biometric solution. However, despite its significant advantages, face…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Arun Kunwar , Ajita Rattani

Foundation models leverage large-scale pretraining to capture extensive knowledge, demonstrating generalization in a wide range of language tasks. By comparison, vision foundation models (VFMs) often exhibit uneven improvements across…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Shiqi Huang , Yipei Wang , Natasha Thorley , Alexander Ng , Shaheer Saeed , Mark Emberton , Shonit Punwani , Veeru Kasivisvanathan , Dean Barratt , Daniel Alexander , Yipeng Hu

Face anti-spoofing (FAS) plays a vital role in preventing face recognition systems from presentation attacks. Existing face anti-spoofing datasets lack diversity due to the insufficient identity and insignificant variance, which limits the…

Computer Vision and Pattern Recognition · Computer Science 2021-12-02 Hangtong Wu , Dan Zen , Yibo Hu , Hailin Shi , Tao Mei

Masked image modeling (MIM) as pre-training is shown to be effective for numerous vision downstream tasks, but how and where MIM works remain unclear. In this paper, we compare MIM with the long-dominant supervised pre-trained models from…

Computer Vision and Pattern Recognition · Computer Science 2022-05-30 Zhenda Xie , Zigang Geng , Jingcheng Hu , Zheng Zhang , Han Hu , Yue Cao

Image-text training like CLIP has dominated the pretraining of vision foundation models in recent years. Subsequent efforts have been made to introduce region-level visual learning into CLIP's pretraining but face scalability challenges due…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Xiaohu Jiang , Yixiao Ge , Yuying Ge , Dachuan Shi , Chun Yuan , Ying Shan

This paper introduces a novel approach to leverage features learned from both supervised and self-supervised paradigms, to improve image classification tasks, specifically for vehicle classification. Two state-of-the-art self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2023-02-02 Shihan Ma , Jidong J. Yang

This paper evaluates DINOv3, a recent large-scale self-supervised vision backbone, for visuomotor diffusion policy learning in robotic manipulation. We investigate whether a purely self-supervised encoder can match or surpass conventional…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 ThankGod Egbe , Peng Wang , Zhihao Guo , Zidong Chen