English
Related papers

Related papers: Do computer vision foundation models learn the low…

200 papers

How well do representations learned by ML models align with those of humans? Here, we consider concept representations learned by deep learning models and evaluate whether they show a fundamental behavioral signature of human concepts, the…

Artificial Intelligence · Computer Science 2024-05-28 Siddhartha K. Vemuri , Raj Sanjay Shah , Sashank Varma

Foundation models (FMs) are large neural networks trained on broad datasets, excelling in downstream tasks with minimal fine-tuning. Human activity recognition in video has advanced with FMs, driven by competition among different…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Thinesh Thiyakesan Ponbagavathi , Kunyu Peng , Alina Roitberg

Masked image modeling (MIM) learns representations with remarkably good fine-tuning performances, overshadowing previous prevalent pre-training approaches such as image classification, instance contrastive learning, and image-text…

Computer Vision and Pattern Recognition · Computer Science 2022-08-25 Yixuan Wei , Han Hu , Zhenda Xie , Zheng Zhang , Yue Cao , Jianmin Bao , Dong Chen , Baining Guo

Despite impressive empirical advances of SSL in solving various tasks, the problem of understanding and characterizing SSL representations learned from input data remains relatively under-explored. We provide a comparative analysis of how…

Computer Vision and Pattern Recognition · Computer Science 2023-06-05 Xavier F. Cadet , Ranya Aloufi , Alain Miranville , Sara Ahmadi-Abhari , Hamed Haddadi

We present DINO-world, a powerful generalist video world model trained to predict future frames in the latent space of DINOv2. By leveraging a pre-trained image encoder and training a future predictor on a large-scale uncurated video…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Federico Baldassarre , Marc Szafraniec , Basile Terver , Vasil Khalidov , Francisco Massa , Yann LeCun , Patrick Labatut , Maximilian Seitzer , Piotr Bojanowski

Video models have recently been applied with success to problems in content generation, novel view synthesis, and, more broadly, world simulation. Many applications in generation and transfer rely on conditioning these models, typically…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Edoardo A. Dominici , Thomas Deixelberger , Konstantinos Vardis , Markus Steinberger

Recent multimodal models such as Contrastive Language-Image Pre-training (CLIP) have shown remarkable ability to align visual and linguistic representations. However, domains where small visual differences carry large semantic significance,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Hiroshi Sasaki

Vision transformers (ViTs) - especially feature foundation models like DINOv2 - learn rich representations useful for many downstream tasks. However, architectural choices (such as positional encoding) can lead to these models displaying…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Moritz Pawlowsky , Antonis Vamvakeros , Alexander Weiss , Anja Bielefeld , Samuel J. Cooper , Ronan Docherty

The vision transformer-based foundation models, such as ViT or Dino-V2, are aimed at solving problems with little or no finetuning of features. Using a setting of prototypical networks, we analyse to what extent such foundation models can…

Computer Vision and Pattern Recognition · Computer Science 2024-02-26 Dmitry Kangin , Plamen Angelov

Contrastive learning has revolutionized the field of computer vision, learning rich representations from unlabeled data, which generalize well to diverse vision tasks. Consequently, it has become increasingly important to explain these…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Fawaz Sammani , Boris Joukovsky , Nikos Deligiannis

An evaluation criterion for safe and trustworthy deep learning is how well the invariances captured by representations of deep neural networks (DNNs) are shared with humans. We identify challenges in measuring these invariances. Prior works…

Computer Vision and Pattern Recognition · Computer Science 2022-12-05 Vedant Nanda , Ayan Majumdar , Camila Kolling , John P. Dickerson , Krishna P. Gummadi , Bradley C. Love , Adrian Weller

Today's computer vision models achieve human or near-human level performance across a wide variety of vision tasks. However, their architectures, data, and learning algorithms differ in numerous ways from those that give rise to human…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Lukas Muttenthaler , Jonas Dippel , Lorenz Linhardt , Robert A. Vandermeulen , Simon Kornblith

We argue that there are many notions of 'similarity' and that models, like humans, should be able to adapt to these dynamically. This contrasts with most representation learning methods, supervised or self-supervised, which learn a fixed…

Computer Vision and Pattern Recognition · Computer Science 2023-06-14 Sagar Vaze , Nicolas Carion , Ishan Misra

Existing text recognition methods usually need large-scale training data. Most of them rely on synthetic training data due to the lack of annotated real images. However, there is a domain gap between the synthetic data and real data, which…

Computer Vision and Pattern Recognition · Computer Science 2023-03-03 Mingkun Yang , Minghui Liao , Pu Lu , Jing Wang , Shenggao Zhu , Hualin Luo , Qi Tian , Xiang Bai

With the rapid improvement of machine learning (ML) models, cognitive scientists are increasingly asking about their alignment with how humans think. Here, we ask this question for computer vision models and human sensitivity to geometric…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Zekun Wang , Sashank Varma

Foundation models are increasingly developed in computational pathology (CPath) given their promise in facilitating many downstream tasks. While recent studies have evaluated task performance across models, less is known about the structure…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Vaibhav Mishra , William Lotter

Art plagiarism detection plays a crucial role in protecting artists' copyrights and intellectual property, yet it remains a challenging problem in forensic analysis. In this paper, we address the task of recognizing plagiarized paintings…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Sophie Zhou , Shu Kong

Machine learning models often struggle with distribution shifts in real-world scenarios, whereas humans exhibit robust adaptation. Models that better align with human perception may achieve higher out-of-distribution generalization. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Mohammad-Javad Darvishi-Bayazi , Md Rifat Arefin , Jocelyn Faubert , Irina Rish

This report analyzes the evolution of key design patterns in computer vision by examining six influential papers. The analysis begins with foundational architectures for image recognition. We review ResNet, which introduced residual…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Radu-Andrei Bourceanu , Neil De La Fuente , Jan Grimm , Andrei Jardan , Andriy Manucharyan , Cornelius Weiss , Daniel Cremers , Roman Pflugfelder

Vision-Language Models (VLMs) are trained on vast amounts of data captured by humans emulating our understanding of the world. However, known as visual illusions, human's perception of reality isn't always faithful to the physical world.…

Artificial Intelligence · Computer Science 2023-11-02 Yichi Zhang , Jiayi Pan , Yuchen Zhou , Rui Pan , Joyce Chai