English
Related papers

Related papers: Do computer vision foundation models learn the low…

200 papers

The human visual system can effortlessly recognize an object under different extrinsic factors such as lighting, object poses, and background, yet current computer vision systems often struggle with these variations. An important step to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Klemen Kotar , Stephen Tian , Hong-Xing Yu , Daniel L. K. Yamins , Jiajun Wu

Foundation models have demonstrated a remarkable ability to learn rich, transferable representations across diverse modalities such as images, text, and audio. In modern machine learning pipelines, these representations often replace raw…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Selene Cerna , Sara Si-Moussi , Wilfried Thuiller , Hadrien Hendrikx , Vincent Miele

Vision foundation models can perform generalized object classification in zero-shot mode, and face/person recognition when they are fine-tuned. However, fine-tuned models suffer from catastrophic forgetting. We create models that perform…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Thomas M Metz , Matthew Q Hill , Alice J O'Toole

Vision-language models (VLMs) excel at broad visual understanding but remain coarse-grained, exhibit visual biases, and miss subtle visual details. Existing training corpora reinforce this limitation by emphasizing general recognition ("Is…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Damiano Marsili , Aditya Mehta , Ryan Y. Lin , Georgia Gkioxari

Visual recognition in low-data regimes requires deep neural networks to learn generalized representations from limited training samples. Recently, CLIP-based methods have shown promising few-shot performance benefited from the contrastive…

Computer Vision and Pattern Recognition · Computer Science 2023-03-06 Renrui Zhang , Xiangfei Hu , Bohao Li , Siyuan Huang , Hanqiu Deng , Hongsheng Li , Yu Qiao , Peng Gao

Large Multimodal Models (LMMs) typically build on ViTs (e.g., CLIP), yet their training with simple random in-batch negatives limits the ability to capture fine-grained visual differences, particularly in geometric scenarios. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Kai Sun , Yushi Bai , Zhen Yang , Jiajie Zhang , Ji Qi , Lei Hou , Juanzi Li

Multiple Object Tracking (MOT) is a computer vision task that has been employed in a variety of sectors. Some common limitations in MOT are varying object appearances, occlusions, or crowded scenes. To address these challenges, machine…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Niels G. Faber , Seyed Sahand Mohammadi Ziabari , Fatemeh Karimi Nejadasl

Common deep neural networks (DNNs) for image classification have been shown to rely on shortcut opportunities (SO) in the form of predictive and easy-to-represent visual factors. This is known as shortcut learning and leads to impaired…

Computer Vision and Pattern Recognition · Computer Science 2021-10-11 Elias Eulig , Piyapat Saranrittichai , Chaithanya Kumar Mummadi , Kilian Rambach , William Beluch , Xiahan Shi , Volker Fischer

We study three intriguing properties of contrastive learning. First, we generalize the standard contrastive loss to a broader family of losses, and we find that various instantiations of the generalized loss perform similarly under the…

Machine Learning · Computer Science 2021-10-26 Ting Chen , Calvin Luo , Lala Li

To handle the large scale of whole slide images in computational pathology, most approaches first tessellate the images into smaller patches, extract features from these patches, and finally aggregate the feature vectors with…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Benedikt Roth , Valentin Koch , Sophia J. Wagner , Julia A. Schnabel , Carsten Marr , Tingying Peng

How do humans learn to acquire a powerful, flexible and robust representation of objects? While much of this process remains unknown, it is clear that humans do not require millions of object labels. Excitingly, recent algorithmic…

Computer Vision and Pattern Recognition · Computer Science 2020-10-19 Robert Geirhos , Kantharaju Narayanappa , Benjamin Mitzkus , Matthias Bethge , Felix A. Wichmann , Wieland Brendel

Visual translation tolerance refers to our capacity to recognize objects over a wide range of different retinal locations. Although translation is perhaps the simplest spatial transform that the visual system needs to cope with, the extent…

Neurons and Cognition · Quantitative Biology 2020-12-09 Ryan Blything , Valerio Biscione , Ivan I. Vankov , Casimir J. H. Ludwig , Jeffrey S. Bowers

Pre-training image representations from the raw text about images enables zero-shot vision transfer to downstream tasks. Through pre-training on millions of samples collected from the internet, multimodal foundation models, such as CLIP,…

Machine Learning · Computer Science 2024-03-18 Chenguang Wang , Ruoxi Jia , Xin Liu , Dawn Song

Foundation models constitute a significant advancement in computer vision: after a single, albeit costly, training phase, they can address a wide array of tasks. In the field of Earth observation, over 75 remote sensing vision foundation…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Pierre Adorni , Minh-Tan Pham , Stéphane May , Sébastien Lefèvre

Foundation models pre-trained on web-scale vision-language data, such as CLIP, are widely used as cornerstones of powerful machine learning systems. While pre-training offers clear advantages for downstream learning, it also endows…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Anjun Hu , Jindong Gu , Francesco Pinto , Konstantinos Kamnitsas , Philip Torr

Deep convolutional neural networks (DCNNs) and the ventral visual pathway share vast architectural and functional similarities in visual challenges such as object recognition. Recent insights have demonstrated that both hierarchical…

Computer Vision and Pattern Recognition · Computer Science 2021-09-22 Leonard E. van Dyck , Roland Kwitt , Sebastian J. Denzler , Walter R. Gruber

Multi-modal foundation models like OpenFlamingo, LLaVA, and GPT-4 are increasingly used for various real-world tasks. Prior work has shown that these models are highly vulnerable to adversarial attacks on the vision modality. These attacks…

Machine Learning · Computer Science 2024-06-06 Christian Schlarmann , Naman Deep Singh , Francesco Croce , Matthias Hein

Biological visual systems learn from limited experience, unlike deep learning models that rely on millions of training images. What learning principles make this possible? We tested whether efficient coding, the idea that neural…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Ananya Passi , Brian S. Robinson , Michael F. Bonner

Self-supervised learning holds the promise of eliminating the need for manual data annotation, enabling models to scale effortlessly to massive datasets and larger architectures. By not being tailored to specific tasks or domains, this…

Retinal image of surrounding objects varies tremendously due to the changes in position, size, pose, illumination condition, background context, occlusion, noise, and nonrigid deformations. But despite these huge variations, our visual…

Computer Vision and Pattern Recognition · Computer Science 2017-02-14 Saeed Reza Kheradpisheh , Mohammad Ganjtabesh , Timothée Masquelier