English
Related papers

Related papers: MM-Retinal: Knowledge-Enhanced Foundational Pretra…

200 papers

Diabetic retinopathy (DR) grading from fundus images has attracted increasing interest in both academic and industrial communities. Most convolutional neural network (CNN) based algorithms treat DR grading as a classification task via…

Computer Vision and Pattern Recognition · Computer Science 2021-03-19 Yehui Yang , Fangxin Shang , Binghong Wu , Dalu Yang , Lei Wang , Yanwu Xu , Wensheng Zhang , Tianzhu Zhang

Multimodal large language models (MLLMs) have shown impressive capabilities, yet they often struggle to effectively capture the fine-grained textual information within images crucial for accurate image translation. This often leads to a…

Computation and Language · Computer Science 2026-04-21 Bo Li , Ningyuan Deng , Tianyu Dong , Shaobo Wang , Shaolin Zhu , Lijie Wen

Fundus images are essential for the early screening and detection of eye diseases. While deep learning models using fundus images have significantly advanced the diagnosis of multiple eye diseases, variations in images from different…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Qian Zeng , Le Zhang , Yipeng Liu , Ce Zhu , Fan Zhang

Vision foundation models (VFMs) are predominantly developed using data-centric methods. These methods require training on vast amounts of data usually with high-quality labels, which poses a bottleneck for most institutions that lack both…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Jiabo Huang , Chen Chen , Lingjuan Lyu

Purpose: Deep learning methods have shown promising results in the segmentation, and detection of diseases in medical images. However, most methods are trained and tested on data from a single source, modality, organ, or disease type,…

Image and Video Processing · Electrical Eng. & Systems 2025-08-20 Nchongmaje Ndipenocha , Alina Mirona , Kezhi Wanga , Yongmin Li

In recent literature, few-shot classification has predominantly been defined by the N-way k-shot meta-learning problem. Models designed for this purpose are usually trained to excel on standard benchmarks following a restricted setup,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Constance Ferragu , Philomene Chagniot , Vincent Coyette

Large-scale public datasets with high-quality annotations are rarely available for intelligent medical imaging research, due to data privacy concerns and the cost of annotations. In this paper, we release SynFundus-1M, a high-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Fangxin Shang , Jie Fu , Yehui Yang , Haifeng Huang , Junwei Liu , Lei Ma

The need for improved diagnostic methods in ophthalmology is acute, especially in the underdeveloped regions with limited access to specialists and advanced equipment. Therefore, we introduce VisionUnite, a novel vision-language foundation…

Image and Video Processing · Electrical Eng. & Systems 2025-08-13 Zihan Li , Diping Song , Zefeng Yang , Deming Wang , Fei Li , Xiulan Zhang , Paul E. Kinahan , Yu Qiao

Fundus photography is prone to suffer from image quality degradation that impacts clinical examination performed by ophthalmologists or intelligent systems. Though enhancement algorithms have been developed to promote fundus observation on…

Image and Video Processing · Electrical Eng. & Systems 2023-09-12 Heng Li , Haofeng Liu , Huazhu Fu , Yanwu Xu , Hui Shu , Ke Niu , Yan Hu , Jiang Liu

Retinal vessel segmentation is generally grounded in image-based datasets collected with bench-top devices. The static images naturally lose the dynamic characteristics of retina fluctuation, resulting in diminished dataset richness, and…

In recent years, deep learning has shown promise in predicting hypertension (HTN) from fundus images. However, most prior research has primarily focused on analyzing a single type of data, which may not capture the full complexity of HTN…

Image and Video Processing · Electrical Eng. & Systems 2024-03-26 Mohammed Baharoon , Hessa Almatar , Reema Alduhayan , Tariq Aldebasi , Badr Alahmadi , Yahya Bokhari , Mohammed Alawad , Ahmed Almazroa , Abdulrhman Aljouie

Prompt tuning, like CoOp, has recently shown promising vision recognizing and transfer learning ability on various downstream tasks with the emergence of large pre-trained vision-language models like CLIP. However, we identify that existing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Yongzhu Miao , Shasha Li , Jintao Tang , Ting Wang

Machine learning is gaining significant attention as a diagnostic tool in medical imaging, particularly in the analysis of retinal fundus images. However, this approach is not yet clinically applicable, as it still depends on human…

Human-Computer Interaction · Computer Science 2025-10-03 Mattea Reid , Zuhairah Zainal , Khaing Zin Than , Danielle Chan , Jonathan Chan

There has been a growing interest in developing multimodal machine translation (MMT) systems that enhance neural machine translation (NMT) with visual knowledge. This problem setup involves using images as auxiliary information during…

Computer Vision and Pattern Recognition · Computer Science 2023-08-30 Devaansh Gupta , Siddhant Kharbanda , Jiawei Zhou , Wanhua Li , Hanspeter Pfister , Donglai Wei

Real-world non-mydriatic retinal fundus photography is prone to artifacts, imperfections and low-quality when certain ocular or systemic co-morbidities exist. Artifacts may result in inaccuracy or ambiguity in clinical diagnoses. In this…

Image and Video Processing · Electrical Eng. & Systems 2023-02-07 Wenhui Zhu , Peijie Qiu , Mohammad Farazi , Keshav Nandakumar , Oana M. Dumitrascu , Yalin Wang

Our research focuses on the critical field of early diagnosis of disease by examining retinal blood vessels in fundus images. While automatic segmentation of retinal blood vessels holds promise for early detection, accurate analysis remains…

Image and Video Processing · Electrical Eng. & Systems 2024-05-14 Fatema Tuj Johora Faria , Mukaffi Bin Moin , Pronay Debnath , Asif Iftekher Fahim , Faisal Muhammad Shah

Diabetic Macular Edema (DME) is a leading cause of vision loss among patients with Diabetic Retinopathy (DR). While deep learning has shown promising results for automatically detecting this condition from fundus images, its application…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Franco Javier Arellano , José Ignacio Orlando

Fine-tuning vision-language models (VLMs) such as CLIP often leads to catastrophic forgetting of pretrained knowledge. Prior work primarily aims to mitigate forgetting during adaptation; however, forgetting often remains inevitable during…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Wenqing Wang , Da Li , Xiatian Zhu , Josef Kittler

Multi-modal retrieval has seen tremendous progress with the development of vision-language models. However, further improving these models require additional labelled data which is a huge manual effort. In this paper, we propose a framework…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Avinash Madasu , Estelle Aflalo , Gabriela Ben Melech Stan , Shachar Rosenman , Shao-Yen Tseng , Gedas Bertasius , Vasudev Lal

Diffusion Transformers (DiTs) have exhibited robust capabilities in image generation tasks. However, accurate text-guided image editing for multimodal DiTs (MM-DiTs) still poses a significant challenge. Unlike UNet-based structures that…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Yu Xu , Fan Tang , Juan Cao , Yuxin Zhang , Xiaoyu Kong , Jintao Li , Oliver Deussen , Tong-Yee Lee