English
Related papers

Related papers: MM-DINOv2: Adapting Foundation Models for Multi-Mo…

200 papers

Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks. However, we observe that their feature representations are poorly aligned across different modalities. For instance, the feature embedding…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Rishabh Kabra , Maks Ovsjanikov , Drew A. Hudson , Ye Xia , Skanda Koppula , Andre Araujo , Joao Carreira , Niloy J. Mitra

In hematology, computational models offer significant potential to improve diagnostic accuracy, streamline workflows, and reduce the tedious work of analyzing single cells in peripheral blood or bone marrow smears. However, clinical…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Valentin Koch , Sophia J. Wagner , Salome Kazeminia , Ece Sancar , Matthias Hehr , Julia Schnabel , Tingying Peng , Carsten Marr

Face Anti-Spoofing (FAS) remains challenging due to the requirement for robust domain generalization across unseen environments. While recent trends leverage Vision-Language Models (VLMs) for semantic supervision, these multimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Mika Feng , Pierre Gallin-Martel , Koichi Ito , Takafumi Aoki

According to the 2021 World Health Organization (WHO) Classification scheme for gliomas, glioma segmentation is a very important basis for diagnosis and genotype prediction. In general, 3D multimodal brain MRI is an effective diagnostic…

Image and Video Processing · Electrical Eng. & Systems 2025-06-23 Xiaoyu Shi , Shurong Chai , Yinhao Li , Jingliang Cheng , Jie Bai , Guohua Zhao , Yen-Wei Chen

Accurate left atrium (LA) segmentation from pre-operative scans is crucial for diagnosing atrial fibrillation, treatment planning, and supporting surgical interventions. While deep learning models are key in medical image segmentation, they…

Image and Video Processing · Electrical Eng. & Systems 2024-11-15 Bipasha Kundu , Bidur Khanal , Richard Simon , Cristian A. Linte

Multi-task image restoration has gained significant interest due to its inherent versatility and efficiency compared to its single-task counterpart. However, performance decline is observed with an increase in the number of tasks, primarily…

Computer Vision and Pattern Recognition · Computer Science 2024-08-19 Xin Lin , Jingtong Yue , Kelvin C. K. Chan , Lu Qi , Chao Ren , Jinshan Pan , Ming-Hsuan Yang

Multimodal learning leverages complementary information derived from different modalities, thereby enhancing performance in medical image segmentation. However, prevailing multimodal learning methods heavily rely on extensive well-annotated…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Xiaogen Zhou , Yiyou Sun , Min Deng , Winnie Chiu Wing Chu , Qi Dou

Medical generative models, acknowledged for their high-quality sample generation ability, have accelerated the fast growth of medical applications. However, recent works concentrate on separate medical generation models for distinct medical…

Image and Video Processing · Electrical Eng. & Systems 2024-03-08 Chenlu Zhan , Yu Lin , Gaoang Wang , Hongwei Wang , Jian Wu

Foundation models (FMs) have emerged as a transformative paradigm in medical image analysis, offering the potential to provide generalizable, task-agnostic solutions across a wide range of clinical tasks and imaging modalities. Their…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Karma Phuntsho , Abdullah , Kyungmi Lee , Ickjai Lee , Euijoon Ahn

Vision-language models, such as CLIP, have achieved significant success in aligning visual and textual representations, becoming essential components of many multi-modal large language models (MLLMs) like LLaVA and OpenFlamingo. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Shizhan Gong , Yankai Jiang , Qi Dou , Farzan Farnia

Recent advances in multimodal foundation models have set new standards in few-shot anomaly detection. This paper explores whether high-quality visual features alone are sufficient to rival existing state-of-the-art vision-language models.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Simon Damm , Mike Laszkiewicz , Johannes Lederer , Asja Fischer

Existing medical image registration algorithms rely on either dataset specific training or local texture-based features to align images. The former cannot be reliably implemented without large modality-specific training datasets, while the…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Xinrui Song , Xuanang Xu , Pingkun Yan

Multimodal medical image fusion (MMIF) aims to integrate images from different modalities to produce a comprehensive image that enhances medical diagnosis by accurately depicting organ structures, tissue textures, and metabolic information.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Tao Luo , Weihua Xu

In this study, we present a multimodal framework for predicting neuro-facial disorders by capturing both vocal and facial cues. We hypothesize that explicitly disentangling shared and modality-specific representations within multimodal…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-13 Mohd Mujtaba Akhtar , Girish , Muskaan Singh

This paper presents a new ambient light normalization framework, DINOLight, that integrates the self-supervised model DINOv2's image understanding capability into the restoration process as a visual prior. Ambient light normalization aims…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Youngjin Oh , Junhyeong Kwon , Nam Ik Cho

Deep Learning approaches in dermatological image classification have shown promising results, yet the field faces significant methodological challenges that impede proper evaluation. This paper presents a dual contribution: first, a…

Image and Video Processing · Electrical Eng. & Systems 2025-02-05 Łukasz Miętkiewicz , Leon Ciechanowski , Dariusz Jemielniak

2D visual foundation models, such as DINOv3, a self-supervised model trained on large-scale natural images, have demonstrated strong zero-shot generalization, capturing both rich global context and fine-grained structural cues. However, an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Yik San Cheng , Runkai Zhao , Weidong Cai

Self-supervised learning holds the promise of eliminating the need for manual data annotation, enabling models to scale effortlessly to massive datasets and larger architectures. By not being tailored to specific tasks or domains, this…

Multimodal pathological images are usually in clinical diagnosis, but computer vision-based multimodal image-assisted diagnosis faces challenges with modality fusion, especially in the absence of expert-annotated data. To achieve the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Qinghua Lin , Guang-Hai Liu , Zuoyong Li , Yang Li , Yuting Jiang , Xiang Wu

In this work, we leverage informative embeddings from foundational models for unsupervised anomaly detection in medical imaging. For small datasets, a memory-bank of normative features can directly be used for anomaly detection which has…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Nico Schulthess , Ender Konukoglu