English
Related papers

Related papers: MM-Retinal: Knowledge-Enhanced Foundational Pretra…

200 papers

The medical imaging community generates a wealth of datasets, many of which are openly accessible and annotated for specific diseases and tasks such as multi-organ or lesion segmentation. Current practices continue to limit model training…

Image and Video Processing · Electrical Eng. & Systems 2024-01-09 Constantin Ulrich , Fabian Isensee , Tassilo Wald , Maximilian Zenk , Michael Baumgartner , Klaus H. Maier-Hein

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Jingzhi Li , Changjiang Luo , Ruoyu Chen , Hua Zhang , Wenqi Ren , Jianhou Gan , Xiaochun Cao

The quality of supervised fine-tuning (SFT) data is crucial for the performance of large multimodal models (LMMs), yet current data enhancement methods often suffer from factual errors and hallucinations due to inadequate visual perception.…

Artificial Intelligence · Computer Science 2025-10-20 Tingqiao Xu , Ziru Zeng , Jiayu Chen

Pretraining with large-scale 3D volumes has a potential for improving the segmentation performance on a target medical image dataset where the training images and annotations are limited. Due to the high cost of acquiring pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2023-06-30 Guotai Wang , Jianghao Wu , Xiangde Luo , Xinglong Liu , Kang Li , Shaoting Zhang

Current video understanding models excel at recognizing "what" is happening but fall short in high-level cognitive tasks like causal reasoning and future prediction, a limitation rooted in their lack of commonsense world knowledge. To…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 L'ea Dubois , Klaus Schmidt , Chengyu Wang , Ji-Hoon Park , Lin Wang , Santiago Munoz

Purpose To develop a computer based method for the automated assessment of image quality in the context of diabetic retinopathy (DR) to guide the photographer. Methods A deep learning framework was trained to grade the images automatically.…

Computer Vision and Pattern Recognition · Computer Science 2017-03-08 Sajib Kumar Saha , Basura Fernando , Jorge Cuadros , Di Xiao , Yogesan Kanagasingam

We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture,…

The development of multi-label deep learning models for retinal disease classification is often hindered by the scarcity of large, expertly annotated clinical datasets due to patient privacy concerns and high costs. The recent release of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Jerry Cao-Xue , Tien Comlekoglu , Keyi Xue , Guanliang Wang , Jiang Li , Gordon Laurie

The quality of a fundus image can be compromised by numerous factors, many of which are challenging to be appropriately and mathematically modeled. In this paper, we introduce a novel diffusion model based framework, named Learning…

Image and Video Processing · Electrical Eng. & Systems 2023-03-09 Puijin Cheng , Li Lin , Yijin Huang , Huaqing He , Wenhan Luo , Xiaoying Tang

General-purpose foundation models have led to recent breakthroughs in artificial intelligence. In remote sensing, self-supervised learning (SSL) and Masked Image Modeling (MIM) have been adopted to build foundation models. However, these…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Fan Liu , Delong Chen , Zhangqingyun Guan , Xiaocong Zhou , Jiale Zhu , Qiaolin Ye , Liyong Fu , Jun Zhou

Existing few-shot segmentation (FSS) methods mainly focus on designing novel support-query matching and self-matching mechanisms to exploit implicit knowledge in pre-trained backbones. However, the performance of these methods is often…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Shijie Chang , Lihe Zhang , Huchuan Lu

Large Vision-Language Models excel at multimodal understanding but struggle to deeply integrate visual information into their predominantly text-based reasoning processes, a key challenge in mirroring human cognition. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Ziwei Zheng , Michael Yang , Jack Hong , Chenxiao Zhao , Guohai Xu , Le Yang , Chao Shen , Xing Yu

Multimodal medical image fusion (MMIF) extracts the most meaningful information from multiple source images, enabling a more comprehensive and accurate diagnosis. Achieving high-quality fusion results requires a careful balance of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Dan He , Weisheng Li , Guofen Wang , Yuping Huang , Shiqiang Liu

Foundation models (FMs) are a popular topic of research in AI. Their ability to generalize to new tasks and datasets without retraining or needing an abundance of data makes them an appealing candidate for applications on specialist…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Marga Don , Stijn Pinson , Blanca Guillen Cebrian , Yuki M. Asano

The performance of diagnostic Computer-Aided Design (CAD) systems for retinal diseases depends on the quality of the retinal images being screened. Thus, many studies have been developed to evaluate and assess the quality of such retinal…

Image and Video Processing · Electrical Eng. & Systems 2024-09-17 Saif Khalid , Hatem A. Rashwan , Saddam Abdulwahab , Mohamed Abdel-Nasser , Facundo Manuel Quiroga , Domenec Puig

With access to large-scale, unlabeled medical datasets, researchers are confronted with two questions: Should they attempt to pretrain a custom foundation model on this medical data, or use transfer-learning from an existing generalist…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Jakob Ambsdorf , Asbjørn Munk , Sebastian Llambias , Anders Nymark Christensen , Kamil Mikolaj , Randall Balestriero , Martin Tolsgaard , Aasa Feragen , Mads Nielsen

Background: RETFound, a self-supervised, retina-specific foundation model (FM), showed potential in downstream applications. However, its comparative performance with traditional deep learning (DL) models remains incompletely understood.…

We identify two major limitations in the existing studies on retinal vessel segmentation: (1) Most existing works are restricted to one modality, i.e., the Color Fundus (CF). However, multi-modality retinal images are used every day in the…

Image and Video Processing · Electrical Eng. & Systems 2026-01-01 Bo Wen , Anna Heinke , Akshay Agnihotri , Dirk-Uwe Bartsch , William Freeman , Truong Nguyen , Cheolhong An

Most existing methods focus on sentiment analysis of textual data. However, recently there has been a massive use of images and videos on social platforms, motivating sentiment analysis from other modalities. Current studies show that…

Machine Learning · Computer Science 2022-10-13 Guilherme Lourenço de Toledo , Ricardo Marcondes Marcacini

Medical image analysis increasingly relies on the integration of multiple imaging modalities to capture complementary anatomical and functional information, enabling more accurate diagnosis and treatment planning. Achieving aligned feature…

Image and Video Processing · Electrical Eng. & Systems 2025-09-04 Yunhao Liu , Suyang Xi , Shiqi Liu , Hong Ding , Chicheng Jin , Chong Zhong , Junjun He , Catherine C. Liu , Yiqing Shen