中文
相关论文

相关论文: MM-Retinal: Knowledge-Enhanced Foundational Pretra…

200 篇论文

Diabetic retinopathy (DR) grading from fundus images has attracted increasing interest in both academic and industrial communities. Most convolutional neural network (CNN) based algorithms treat DR grading as a classification task via…

计算机视觉与模式识别 · 计算机科学 2021-03-19 Yehui Yang , Fangxin Shang , Binghong Wu , Dalu Yang , Lei Wang , Yanwu Xu , Wensheng Zhang , Tianzhu Zhang

Multimodal large language models (MLLMs) have shown impressive capabilities, yet they often struggle to effectively capture the fine-grained textual information within images crucial for accurate image translation. This often leads to a…

计算与语言 · 计算机科学 2026-04-21 Bo Li , Ningyuan Deng , Tianyu Dong , Shaobo Wang , Shaolin Zhu , Lijie Wen

Fundus images are essential for the early screening and detection of eye diseases. While deep learning models using fundus images have significantly advanced the diagnosis of multiple eye diseases, variations in images from different…

计算机视觉与模式识别 · 计算机科学 2025-11-10 Qian Zeng , Le Zhang , Yipeng Liu , Ce Zhu , Fan Zhang

Vision foundation models (VFMs) are predominantly developed using data-centric methods. These methods require training on vast amounts of data usually with high-quality labels, which poses a bottleneck for most institutions that lack both…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Jiabo Huang , Chen Chen , Lingjuan Lyu

Purpose: Deep learning methods have shown promising results in the segmentation, and detection of diseases in medical images. However, most methods are trained and tested on data from a single source, modality, organ, or disease type,…

图像与视频处理 · 电气工程与系统科学 2025-08-20 Nchongmaje Ndipenocha , Alina Mirona , Kezhi Wanga , Yongmin Li

In recent literature, few-shot classification has predominantly been defined by the N-way k-shot meta-learning problem. Models designed for this purpose are usually trained to excel on standard benchmarks following a restricted setup,…

计算机视觉与模式识别 · 计算机科学 2024-05-21 Constance Ferragu , Philomene Chagniot , Vincent Coyette

Large-scale public datasets with high-quality annotations are rarely available for intelligent medical imaging research, due to data privacy concerns and the cost of annotations. In this paper, we release SynFundus-1M, a high-quality…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Fangxin Shang , Jie Fu , Yehui Yang , Haifeng Huang , Junwei Liu , Lei Ma

The need for improved diagnostic methods in ophthalmology is acute, especially in the underdeveloped regions with limited access to specialists and advanced equipment. Therefore, we introduce VisionUnite, a novel vision-language foundation…

图像与视频处理 · 电气工程与系统科学 2025-08-13 Zihan Li , Diping Song , Zefeng Yang , Deming Wang , Fei Li , Xiulan Zhang , Paul E. Kinahan , Yu Qiao

Fundus photography is prone to suffer from image quality degradation that impacts clinical examination performed by ophthalmologists or intelligent systems. Though enhancement algorithms have been developed to promote fundus observation on…

图像与视频处理 · 电气工程与系统科学 2023-09-12 Heng Li , Haofeng Liu , Huazhu Fu , Yanwu Xu , Hui Shu , Ke Niu , Yan Hu , Jiang Liu

Retinal vessel segmentation is generally grounded in image-based datasets collected with bench-top devices. The static images naturally lose the dynamic characteristics of retina fluctuation, resulting in diminished dataset richness, and…

In recent years, deep learning has shown promise in predicting hypertension (HTN) from fundus images. However, most prior research has primarily focused on analyzing a single type of data, which may not capture the full complexity of HTN…

图像与视频处理 · 电气工程与系统科学 2024-03-26 Mohammed Baharoon , Hessa Almatar , Reema Alduhayan , Tariq Aldebasi , Badr Alahmadi , Yahya Bokhari , Mohammed Alawad , Ahmed Almazroa , Abdulrhman Aljouie

Prompt tuning, like CoOp, has recently shown promising vision recognizing and transfer learning ability on various downstream tasks with the emergence of large pre-trained vision-language models like CLIP. However, we identify that existing…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Yongzhu Miao , Shasha Li , Jintao Tang , Ting Wang

Machine learning is gaining significant attention as a diagnostic tool in medical imaging, particularly in the analysis of retinal fundus images. However, this approach is not yet clinically applicable, as it still depends on human…

人机交互 · 计算机科学 2025-10-03 Mattea Reid , Zuhairah Zainal , Khaing Zin Than , Danielle Chan , Jonathan Chan

There has been a growing interest in developing multimodal machine translation (MMT) systems that enhance neural machine translation (NMT) with visual knowledge. This problem setup involves using images as auxiliary information during…

计算机视觉与模式识别 · 计算机科学 2023-08-30 Devaansh Gupta , Siddhant Kharbanda , Jiawei Zhou , Wanhua Li , Hanspeter Pfister , Donglai Wei

Real-world non-mydriatic retinal fundus photography is prone to artifacts, imperfections and low-quality when certain ocular or systemic co-morbidities exist. Artifacts may result in inaccuracy or ambiguity in clinical diagnoses. In this…

图像与视频处理 · 电气工程与系统科学 2023-02-07 Wenhui Zhu , Peijie Qiu , Mohammad Farazi , Keshav Nandakumar , Oana M. Dumitrascu , Yalin Wang

Our research focuses on the critical field of early diagnosis of disease by examining retinal blood vessels in fundus images. While automatic segmentation of retinal blood vessels holds promise for early detection, accurate analysis remains…

图像与视频处理 · 电气工程与系统科学 2024-05-14 Fatema Tuj Johora Faria , Mukaffi Bin Moin , Pronay Debnath , Asif Iftekher Fahim , Faisal Muhammad Shah

Diabetic Macular Edema (DME) is a leading cause of vision loss among patients with Diabetic Retinopathy (DR). While deep learning has shown promising results for automatically detecting this condition from fundus images, its application…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Franco Javier Arellano , José Ignacio Orlando

Fine-tuning vision-language models (VLMs) such as CLIP often leads to catastrophic forgetting of pretrained knowledge. Prior work primarily aims to mitigate forgetting during adaptation; however, forgetting often remains inevitable during…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Wenqing Wang , Da Li , Xiatian Zhu , Josef Kittler

Multi-modal retrieval has seen tremendous progress with the development of vision-language models. However, further improving these models require additional labelled data which is a huge manual effort. In this paper, we propose a framework…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Avinash Madasu , Estelle Aflalo , Gabriela Ben Melech Stan , Shachar Rosenman , Shao-Yen Tseng , Gedas Bertasius , Vasudev Lal

Diffusion Transformers (DiTs) have exhibited robust capabilities in image generation tasks. However, accurate text-guided image editing for multimodal DiTs (MM-DiTs) still poses a significant challenge. Unlike UNet-based structures that…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Yu Xu , Fan Tang , Juan Cao , Yuxin Zhang , Xiaoyu Kong , Jintao Li , Oliver Deussen , Tong-Yee Lee