English
Related papers

Related papers: Zero-Ablation Overstates Register Content Dependen…

200 papers

We investigate the mechanism underlying a previously identified phenomenon in Vision Transformers - the emergence of high-norm tokens that lead to noisy attention maps (Darcet et al., 2024). We observe that in multiple models (e.g., CLIP,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Nick Jiang , Amil Dravid , Alexei Efros , Yossi Gandelsman

Training Vision Transformers (ViTs) presents significant challenges, one of which is the emergence of artifacts in attention maps, hindering their interpretability. Darcet et al. (2024) investigated this phenomenon and attributed it to the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Spiros Baxevanakis , Platon Karageorgis , Ioannis Dravilas , Konrad Szewczyk

Vision transformers (ViTs) - especially feature foundation models like DINOv2 - learn rich representations useful for many downstream tasks. However, architectural choices (such as positional encoding) can lead to these models displaying…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Moritz Pawlowsky , Antonis Vamvakeros , Alexander Weiss , Anja Bielefeld , Samuel J. Cooper , Ronan Docherty

Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a problem effectively mitigated by \textit{register tokens}. As diffusion models increasingly adopt transformer architectures…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Nikita Starodubcev , Ilia Sudakov , Ilya Drobyshevskiy , Artem Babenko , Dmitry Baranchuk

Vision Transformers (ViTs) have shown success across a variety of tasks due to their ability to capture global image representations. Recent studies have identified the existence of high-norm tokens in ViTs, which can interfere with…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Srikar Yellapragada , Kowshik Thopalli , Vivek Narayanaswamy , Wesam Sakla , Yang Liu , Yamen Mubarka , Dimitris Samaras , Jayaraman J. Thiagarajan

Transformer-based models have dominated natural language processing and other areas in the last few years due to their superior (zero-shot) performance on benchmark datasets. However, these models are poorly understood due to their…

Computer Vision and Pattern Recognition · Computer Science 2024-02-14 Shaeke Salman , Md Montasir Bin Shams , Xiuwen Liu , Lingjiong Zhu

Drawing inspiration from recent findings including surprisingly decent performance of transformers without positional encoding (NoPE) in the domain of language models and how registers (additional throwaway tokens not tied to input) may…

Computation and Language · Computer Science 2026-01-23 Jason Chuan-Chih Chou , Abhinav Kumar , Shivank Garg

The use of transformer-based models is growing rapidly throughout society. With this growth, it is important to understand how they work, and in particular, how the attention mechanisms represent concepts. Though there are many…

Machine Learning · Computer Science 2024-09-02 Nicholas Pochinkov , Ben Pasero , Skylar Shibayama

We address zero-shot TTS systems' noise-robustness problem by proposing a dual-objective training for the speaker encoder using self-supervised DINO loss. This approach enhances the speaker encoder with the speech synthesis objective,…

Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding tasks. However, the increasing demand for high-resolution image and long-video understanding results in substantial token counts,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Junjie Chen , Xuyang Liu , Zichen Wen , Yiyu Wang , Siteng Huang , Honggang Chen

Recently vision transformers have been shown to be competitive with convolution-based methods (CNNs) broadly across multiple vision tasks. The less restrictive inductive bias of transformers endows greater representational capacity in…

Computer Vision and Pattern Recognition · Computer Science 2022-10-27 Farrukh Rahman , Ömer Mubarek , Zsolt Kira

DINO and DINOv2 are two model families being widely used to learn representations from unlabeled imagery data at large scales. Their learned representations often enable state-of-the-art performance for downstream tasks, such as image…

Computer Vision and Pattern Recognition · Computer Science 2025-02-17 Ziyang Wu , Jingyuan Zhang , Druv Pai , XuDong Wang , Chandan Singh , Jianwei Yang , Jianfeng Gao , Yi Ma

The design, optimisation and construction of an anti-coincidence veto detector to complement the ZEPLIN-III direct dark matter search instrument is described. One tonne of plastic scintillator is arranged into 52 bars individually read out…

Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models such as CLIP, self-supervised visual features are not…

Vision Transformers (ViTs) dominate self-supervised learning (SSL). While they have proven highly effective for large-scale pretraining, they are computationally inefficient and scale poorly with image size. Consequently, foundational…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Nedyalko Prisadnikov , Danda Pani Paudel , Yuqian Fu , Luc Van Gool

Vision Transformers (ViTs) have demonstrated superior performance over Convolutional Neural Networks (CNNs) in various vision-related tasks such as classification, object detection, and segmentation due to their use of self-attention…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Fereshteh Baradaran , Mohsen Raji , Azadeh Baradaran , Arezoo Baradaran , Reihaneh Akbarifard

Vision-language models trained on large, randomly collected data had significant impact in many areas since they appeared. But as they show great performance in various fields, such as image-text-retrieval, their inner workings are still…

Computer Vision and Pattern Recognition · Computer Science 2022-09-15 Felix Vogel , Nina Shvetsova , Leonid Karlinsky , Hilde Kuehne

Recent work has shown that the attention maps of the widely popular DINOv2 model exhibit artifacts, which hurt both model interpretability and performance on dense image tasks. These artifacts emerge due to the model repurposing patch…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Alexander Lappe , Martin A. Giese

We propose to learn invariant representations, in the data domain, to achieve interpretability in algorithmic fairness. Invariance implies a selectivity for high level, relevant correlations w.r.t. class label annotations, and a robustness…

Machine Learning · Computer Science 2020-08-13 Thomas Kehrenberg , Myles Bartlett , Oliver Thomas , Novi Quadrianto

Adapting foundation models to medical segmentation typically requires either backbone fine-tuning or high-capacity task-specific decoders, both of which are difficult to fit reliably when annotations are scarce. We show that frozen DINOv3…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Wei Jiang , Feng Liu , Nan Ye , Hongfu Sun
‹ Prev 1 2 3 10 Next ›