English
Related papers

Related papers: U-REPA: Aligning Diffusion U-Nets to ViTs

200 papers

Most methods for medical image segmentation use U-Net or its variants as they have been successful in most of the applications. After a detailed analysis of these "traditional" encoder-decoder based approaches, we observed that they perform…

Image and Video Processing · Electrical Eng. & Systems 2021-10-18 Jeya Maria Jose Valanarasu , Vishwanath A. Sindagi , Ilker Hacihaliloglu , Vishal M. Patel

Vision Transformer (ViT) has gained increasing attention in the computer vision community in recent years. However, the core component of ViT, Self-Attention, lacks explicit spatial priors and bears a quadratic computational complexity,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Qihang Fan , Huaibo Huang , Mingrui Chen , Hongmin Liu , Ran He

U-Nets have been established as a standard architecture for image-to-image learning problems such as segmentation and inverse problems in imaging. For large-scale data, as it for example appears in 3D medical imaging, the U-Net however has…

Machine Learning · Computer Science 2020-07-01 Christian Etmann , Rihuan Ke , Carola-Bibiane Schönlieb

Unsupervised Domain Adaptation (UDA) aims to enhance the generalization of the learned model to other domains. The domain-invariant knowledge is transferred from the model trained on labeled source domain, e.g., video game, to unlabeled…

Computer Vision and Pattern Recognition · Computer Science 2022-11-15 Mu Chen , Zhedong Zheng , Yi Yang , Tat-Seng Chua

Uncertainty estimation in machine learning is paramount for enhancing the reliability and interpretability of predictive models, especially in high-stakes real-world scenarios. Despite the availability of numerous methods, they often pose a…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Anton Baumann , Thomas Roßberg , Michael Schmitt

The manifold hypothesis posits that high-dimensional data often lies on a lower-dimensional manifold and that utilizing this manifold as the target space yields more efficient representations. While numerous traditional manifold-based…

Machine Learning · Computer Science 2024-06-25 Li Meng , Morten Goodwin , Anis Yazidi , Paal Engelstad

In this paper, we propose a new adapter network for adapting a pre-trained deep neural network to a target domain with minimal computation. The proposed model, unidirectional thin adapter (UDTA), helps the classifier adapt to new data by…

Computer Vision and Pattern Recognition · Computer Science 2022-03-24 Han Gyel Sun , Hyunjae Ahn , HyunGyu Lee , Injung Kim

Reference-to-video (R2V) generation is a controllable video synthesis paradigm that constrains the generation process using both text prompts and reference images, enabling applications such as personalized advertising and virtual try-on.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Lei Wang , YuXin Song , Ge Wu , Haocheng Feng , Hang Zhou , Jingdong Wang , Yaxing Wang , jian Yang

Deep unsupervised domain adaptation (UDA) has recently received increasing attention from researchers. However, existing methods are computationally intensive due to the computation cost of Convolutional Neural Networks (CNN) adopted by…

Machine Learning · Computer Science 2019-04-05 Chaohui Yu , Jindong Wang , Yiqiang Chen , Zijing Wu

Most face identification approaches employ a Siamese neural network to compare two images at the image embedding level. Yet, this technique can be subject to occlusion (e.g. faces with masks or sunglasses) and out-of-distribution data.…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Hai Phan , Cindy Le , Vu Le , Yihui He , Anh Totti Nguyen

We present V-JEPA 2.1, a family of self-supervised models that learn dense, high-quality visual representations for both images and videos while retaining strong global scene understanding. The approach combines four key components. First,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Lorenzo Mur-Labadia , Matthew Muckley , Amir Bar , Mido Assran , Koustuv Sinha , Mike Rabbat , Yann LeCun , Nicolas Ballas , Adrien Bardes

The retroperitoneum hosts a variety of tumors, including rare benign and malignant types, which pose diagnostic and treatment challenges due to their infrequency and proximity to vital structures. Estimating tumor volume is difficult due to…

Image and Video Processing · Electrical Eng. & Systems 2025-02-04 Moein Heidari , Ehsan Khodapanah Aghdam , Alexander Manzella , Daniel Hsu , Rebecca Scalabrino , Wenjin Chen , David J. Foran , Ilker Hacihaliloglu

Automatic segmentation of retinal vessels in fundus images plays an important role in the diagnosis of some diseases such as diabetes and hypertension. In this paper, we propose Deformable U-Net (DUNet), which exploits the retinal vessels'…

Computer Vision and Pattern Recognition · Computer Science 2019-05-21 Qiangguo Jin , Zhaopeng Meng , Tuan D. Pham , Qi Chen , Leyi Wei , Ran Su

Background: Underwater images, in general, suffer from low contrast and high color distortions due to the non-uniform attenuation of the light as it propagates through the water. In addition, the degree of attenuation varies with the…

Image and Video Processing · Electrical Eng. & Systems 2022-01-20 Prasen Kumar Sharma , Ira Bisht , Arijit Sur

Vision Transformers (ViTs) incur significant computational overhead due to the quadratic complexity of self-attention relative to the token sequence length. While existing token reduction methods mitigate this issue, they predominantly rely…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Kaixuan He , Song Chen , Yi Kang

Direct image-to-image alignment that relies on the optimization of photometric error metrics suffers from limited convergence range and sensitivity to lighting conditions. Deep learning approaches has been applied to address this problem by…

Computer Vision and Pattern Recognition · Computer Science 2018-12-27 Lei Han , Mengqi Ji , Lu Fang , Matthias Nießner

A multitude of imaging and vision tasks have seen recently a major transformation by deep learning methods and in particular by the application of convolutional neural networks. These methods achieve impressive results, even for…

Computer Vision and Pattern Recognition · Computer Science 2018-11-30 Simon Arridge , Andreas Hauptmann

We introduce a two-stage self-supervised framework that combines the Joint-Embedding Predictive Architecture (JEPA) with a Density Adaptive Attention Mechanism (DAAM) for learning robust speech representations. Stage~1 uses JEPA with DAAM…

U-Net style networks are commonly utilized in unsupervised image registration to predict dense displacement fields, which for high-resolution volumetric image data is a resource-intensive and time-consuming task. To tackle this challenge,…

Image and Video Processing · Electrical Eng. & Systems 2023-07-07 Xi Jia , Alexander Thorley , Alberto Gomez , Wenqi Lu , Dipak Kotecha , Jinming Duan

Point tracking aims to localize corresponding points across video frames, serving as a fundamental task for 4D reconstruction, robotics, and video editing. Existing methods commonly rely on shallow convolutional backbones such as ResNet…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Soowon Son , Honggyu An , Chaehyun Kim , Hyunah Ko , Jisu Nam , Dahyun Chung , Siyoon Jin , Jung Yi , Jaewon Min , Junhwa Hur , Seungryong Kim