English
Related papers

Related papers: MM-DINOv2: Adapting Foundation Models for Multi-Mo…

200 papers

The data-intensive nature of supervised classification drives the interest of the researchers towards unsupervised approaches, especially for problems such as medical image segmentation, where labeled data is scarce. Building on the recent…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 A. Mudit Adityaja , Saurabh J. Shigwan , Nitin Kumar

Convolutional networks, transformers, hybrid models, and Mamba-based architectures have demonstrated strong performance across various medical image classification tasks. However, these methods were primarily designed to classify clean…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Omid Nejati Manzari , Hojat Asgariandehkordi , Taha Koleilat , Yiming Xiao , Hassan Rivaz

Atypical mitotic figures (AMFs) represent abnormal cell division associated with poor prognosis. Yet their detection remains difficult due to low prevalence, subtle morphology, and inter-observer variability. The MIDOG 2025 challenge…

Image and Video Processing · Electrical Eng. & Systems 2025-10-15 Guillaume Balezo , Hana Feki , Raphaël Bourgade , Lily Monnier , Matthieu Blons , Alice Blondel , Etienne Decencière , Albert Pla Planas , Thomas Walter

Fluoroscopy is critical for real-time X-ray visualization in medical imaging. However, low-dose images are compromised by noise, potentially affecting diagnostic accuracy. Noise reduction is crucial for maintaining image quality, especially…

Image and Video Processing · Electrical Eng. & Systems 2024-11-05 Sun-Young Jeon , Sen Wang , Adam S. Wang , Garry E. Gold , Jang-Hwan Choi

T2 hyperintensities in spinal cord MR images are crucial biomarkers for conditions such as degenerative cervical myelopathy. However, current clinical diagnoses primarily rely on manual evaluation. Deep learning methods have shown promise…

Image and Video Processing · Electrical Eng. & Systems 2025-03-18 Qi Zhang , Xiuyuan Chen , Ziyi He , Kun Wang , Lianming Wu , Hongxing Shen , Jianqi Sun

Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification, and segmentation. However, such representations tend to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Jona Ruthardt , Manu Gaur , Deva Ramanan , Makarand Tapaswi , Yuki M. Asano

We present a novel method for scene change detection that leverages the robust feature extraction capabilities of a visual foundational model, DINOv2, and integrates full-image cross-attention to address key challenges such as varying…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Chun-Jung Lin , Sourav Garg , Tat-Jun Chin , Feras Dayoub

Vision Transformers (ViTs), such as DINOv2, achieve strong performance across domains but often repurpose low-informative patch tokens in ways that reduce the interpretability of attention and feature maps. This challenge is especially…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Joel Valdivia Ortega , Lorenz Lamm , Franziska Eckardt , Benedikt Schworm , Marion Jasnin , Tingying Peng

Recent studies on contrastive learning have achieved remarkable performance solely by leveraging few labels in the context of medical image segmentation. Existing methods mainly focus on instance discrimination and invariant mapping.…

Image and Video Processing · Electrical Eng. & Systems 2024-09-24 Chenyu You , Weicheng Dai , Fenglin Liu , Yifei Min , Nicha C. Dvornek , Xiaoxiao Li , David A. Clifton , Lawrence Staib , James S. Duncan

In this work we investigate the viability of foundational AI/ML models for Synthetic Aperture Radar (SAR) object recognition tasks. We are inspired by the tremendous progress being made in the wider community, particularly in the natural…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Nathan Inkawhich

We present a multi-modal classification framework that fuses satellite and street-level imagery through a Perceiver IO architecture operating on spatial patch tokens from a shared DINOv2 backbone. The design naturally handles a variable…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Niels Sombekke , Rob G. J. Wijnhoven , Martin R. Oswald

Anomaly detection in medical images is an important yet challenging task due to the diversity of possible anomalies and the practical impossibility of collecting comprehensively annotated data sets. In this work, we tackle unsupervised…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Francesco Dalmonte , Emirhan Bayar , Emre Akbas , Mariana-Iuliana Georgescu

Incomplete multi-modal medical image segmentation faces critical challenges from modality imbalance, including imbalanced modality missing rates and heterogeneous modality contributions. Due to their reliance on idealized assumptions of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Libin Lan , Hongxing Li , Zunhui Xia , Yudong Zhang

Birds Eye View perception models require extensive data to perform and generalize effectively. While traditional datasets often provide abundant driving scenes from diverse locations, this is not always the case. It is crucial to maximize…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Seamie Hayes , Ganesh Sistu , Ciarán Eising

Retinal diseases spanning a broad spectrum can be effectively identified and diagnosed using complementary signals from multimodal data. However, multimodal diagnosis in ophthalmic practice is typically challenged in terms of data…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Lu Zhang , Huizhen Yu , Zuowei Wang , Fu Gui , Yatu Guo , Wei Zhang , Mengyu Jia

Brain MRI underpins a wide range of neuroscientific and clinical applications, yet most learning-based methods remain task-specific and require substantial labeled data. Here we show that a single self-supervised representation can…

Machine Learning · Computer Science 2026-05-27 Yizhou Wu , Shansong Wang , Yuheng Li , Mojtaba Safari , Mingzhe Hu , Chih-Wei Chang , Harini Veeraraghavan , Xiaofeng Yang

Recently, Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in visual understanding and reasoning across various vision-language tasks. However, we found that MLLMs cannot process effectively from…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Bangyan Li , Wenxuan Huang , Zhenkun Gao , Yeqiang Wang , Yunhang Shen , Jingzhong Lin , Ling You , Yuxiang Shen , Shaohui Lin , Wanli Ouyang , Yuling Sun

The availability of handy multi-modal (i.e., RGB-D) sensors has brought about a surge of face anti-spoofing research. However, the current multi-modal face presentation attack detection (PAD) has two defects: (1) The framework based on…

Computer Vision and Pattern Recognition · Computer Science 2023-05-08 Ajian Liu , Zichang Tan , Zitong Yu , Chenxu Zhao , Jun Wan , Yanyan Liang , Zhen Lei , Du Zhang , Stan Z. Li , Guodong Guo

The ability to predict future outcomes given control actions is fundamental for physical reasoning. However, such predictive models, often called world models, remains challenging to learn and are typically developed for task-specific…

Robotics · Computer Science 2025-02-04 Gaoyue Zhou , Hengkai Pan , Yann LeCun , Lerrel Pinto

Deep learning has shown significant value in medical image registration for motion correction, however, current techniques are either limited by the type and range of motion they can handle, or require iterative inference and/or retraining…

Image and Video Processing · Electrical Eng. & Systems 2026-05-04 Jian Wang , Razieh Faghihpirayesh , Danny Joca , Polina Golland , Ali Gholipour
‹ Prev 1 8 9 10 Next ›