English
Related papers

Related papers: DinoLizer: Learning from the Best for Generative I…

200 papers

Recent advances in diffusion models (DMs) have achieved exceptional visual quality in image editing tasks. However, the global denoising dynamics of DMs inherently conflate local editing targets with the full-image context, leading to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Wei Chow , Linfeng Li , Lingdong Kong , Zefeng Li , Qi Xu , Hang Song , Tian Ye , Xian Wang , Jinbin Bai , Shilin Xu , Xiangtai Li , Junting Pan , Shaoteng Liu , Ran Zhou , Tianshu Yang , Songhua Liu

Self-supervised learning has proven to be invaluable in making best use of all of the available data in biomedical image segmentation. One particularly simple and effective mechanism to achieve self-supervision is inpainting, the task of…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Subhradeep Kayal , Shuai Chen , Marleen de Bruijne

In this paper, we explore a new generative approach for learning visual representations. Our method, DARL, employs a decoder-only Transformer to predict image patches autoregressively. We find that training with Mean Squared Error (MSE)…

Machine Learning · Computer Science 2024-06-05 Yazhe Li , Jorg Bornschein , Ting Chen

Computer vision methods that explicitly detect object parts and reason on them are a step towards inherently interpretable models. Existing approaches that perform part discovery driven by a fine-grained classification task make very…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Ananthu Aniraj , Cassio F. Dantas , Dino Ienco , Diego Marcos

Recent advances in diffusion models have significantly improved the performance of reference-guided line art colorization. However, existing methods still struggle with region-level color consistency, especially when the reference and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Qianru Qiu , Jiafeng Mao , Kento Masui , Xueting Wang

2D visual foundation models, such as DINOv3, a self-supervised model trained on large-scale natural images, have demonstrated strong zero-shot generalization, capturing both rich global context and fine-grained structural cues. However, an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Yik San Cheng , Runkai Zhao , Weidong Cai

Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent object segmentation properties. However, their alignment with human object perception remains poorly…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Hossein Adeli , Seoyoung Ahn , Andrew Luo , Mengmi Zhang , Nikolaus Kriegeskorte , Gregory Zelinsky

Understanding model decisions is crucial in medical imaging, where interpretability directly impacts clinical trust and adoption. Vision Transformers (ViTs) have demonstrated state-of-the-art performance in diagnostic imaging; however,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Leili Barekatain , Ben Glocker

Most deep learning based image inpainting approaches adopt autoencoder or its variants to fill missing regions in images. Encoders are usually utilized to learn powerful representational spaces, which are important for dealing with…

Computer Vision and Pattern Recognition · Computer Science 2020-10-30 Xin Ma , Xiaoqiang Zhou , Huaibo Huang , Zhenhua Chai , Xiaolin Wei , Ran He

Significant advancements in the field of wood species identification are needed worldwide to support sustainable timber trade. In this work we contribute to automate the identification of wood species via high-resolution macroscopic images…

Anomaly detection and classification in medical imaging are critical for early diagnosis but remain challenging due to limited annotated data, class imbalance, and the high cost of expert labeling. Emerging vision foundation models such as…

Image and Video Processing · Electrical Eng. & Systems 2025-09-17 Fazle Rafsani , Jay Shah , Catherine D. Chong , Todd J. Schwedt , Teresa Wu

This study proposes a retinal prosthetic simulation framework driven by visual fixations, inspired by the saccade mechanism, and assesses performance improvements through end-to-end optimization in a classification task. Salient patches are…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Yuli Wu , Do Dinh Tan Nguyen , Henning Konermann , Rüveyda Yilmaz , Peter Walter , Johannes Stegmaier

In recent years, self-supervised denoising methods have shown impressive performance, which circumvent painstaking collection procedure of noisy-clean image pairs in supervised denoising methods and boost denoising applicability in real…

Image and Video Processing · Electrical Eng. & Systems 2021-09-13 Yuhongze Zhou , Liguang Zhou , Tin Lun Lam , Yangsheng Xu

Images captured from the real world are often affected by different types of noise, which can significantly impact the performance of Computer Vision systems and the quality of visual data. This study presents a novel approach for defect…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Mohsen Hami , Mahdi JameBozorg

Vision Transformers rely on fixed patch tokens that ignore the spatial and semantic structure of images. In this work, we introduce an end-to-end differentiable tokenizer that adapts to image content with pixel-level granularity while…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Marius Aasan , Martine Hjelkrem-Tan , Nico Catalano , Changkyu Choi , Adín Ramírez Rivera

We introduce DenoMAE2.0, an enhanced denoising masked autoencoder that integrates a local patch classification objective alongside traditional reconstruction loss to improve representation learning and robustness. Unlike conventional Masked…

Machine Learning · Computer Science 2025-02-26 Atik Faysal , Mohammad Rostami , Taha Boushine , Reihaneh Gh. Roshan , Huaxia Wang , Nikhil Muralidhar

An accurate and robust large-scale localization system is an integral component for active areas of research such as autonomous vehicles and augmented reality. To this end, many learning algorithms have been proposed that predict 6DOF…

Computer Vision and Pattern Recognition · Computer Science 2022-03-02 Ali Raza , Lazar Lolic , Shahmir Akhter , Alfonso Dela Cruz , Michael Liut

Diffusion models have emerged as powerful priors for image editing tasks such as inpainting and local modification, where the objective is to generate realistic content that remains consistent with observed regions. In particular, zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Badr Moufad , Navid Bagheri Shouraki , Alain Oliviero Durmus , Thomas Hirtz , Eric Moulines , Jimmy Olsson , Yazid Janati

Many models of visual attention have been proposed so far. Traditional bottom-up models, like saliency models, fail to replicate human gaze patterns, and deep gaze prediction models lack biological plausibility due to their reliance on…

Neurons and Cognition · Quantitative Biology 2025-05-28 Takuto Yamamoto , Hirosato Akahoshi , Shigeru Kitazawa

Patch attacks, one of the most threatening forms of physical attack in adversarial examples, can lead networks to induce misclassification by modifying pixels arbitrarily in a continuous region. Certifiable patch defense can guarantee…

Computer Vision and Pattern Recognition · Computer Science 2022-03-17 Zhaoyu Chen , Bo Li , Jianghe Xu , Shuang Wu , Shouhong Ding , Wenqiang Zhang