中文
相关论文

相关论文: MM-Retinal: Knowledge-Enhanced Foundational Pretra…

200 篇论文

Vision-language pretraining (VLP) has been investigated to generalize across diverse downstream tasks for fundus image analysis. Although recent methods showcase promising achievements, they significantly rely on large-scale private…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Ruiqi Wu , Na Su , Chenran Zhang , Tengfei Ma , Tao Zhou , Zhiting Cui , Nianfeng Tang , Tianyu Mao , Yi Zhou , Wen Fan , Tianxing Wu , Shenqi Jing , Huazhu Fu

Traditional fundus image analysis models focus on single-modal tasks, ignoring fundus modality complementarity, which limits their versatility. Recently, retinal foundation models have emerged, but most still remain modality-specific.…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Yuang Yao , Ruiqi Wu , Yi Zhou , Tao Zhou

Artificial intelligence applied to retinal images offers significant potential for recognizing signs and symptoms of retinal conditions and expediting the diagnosis of eye diseases and systemic disorders. However, developing generalized…

图像与视频处理 · 电气工程与系统科学 2024-08-19 Boa Jang , Youngbin Ahn , Eun Kyung Choe , Chang Ki Yoon , Hyuk Jin Choi , Young-Gon Kim

Fundus imaging such as CFP, OCT and UWF is crucial for the early detection of retinal anomalies and diseases. Fundus image understanding, due to its knowledge-intensive nature, poses a challenging vision-language task. An emerging approach…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Yuchuan Deng , Qijie Wei , Kaiheng Qian , Jiazhen Liu , Zijie Xin , Bangxiang Lan , Jingyu Liu , Jianfeng Dong , Xirong Li

Large-scale contrastive vision-language pre-trained models provide the zero-shot model achieving competitive performance across a range of image classification tasks without requiring training on downstream data. Recent works have confirmed…

机器学习 · 计算机科学 2024-04-02 Giung Nam , Byeongho Heo , Juho Lee

Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL). However, many of these models rely on architectures that offer limited interpretability,…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Samuel Ofosu Mensah , Camila Roa , Kerol Djoumessi , Philipp Berens

Retinal fundus photography offers a non-invasive way to diagnose and monitor a variety of retinal diseases, but is prone to inherent quality glitches arising from systemic imperfections or operator/patient-related factors. However,…

图像与视频处理 · 电气工程与系统科学 2024-09-13 Vamsi Krishna Vasa , Peijie Qiu , Wenhui Zhu , Yujian Xiong , Oana Dumitrascu , Yalin Wang

Retinal foundation models aim to learn generalizable representations from diverse retinal images, facilitating label-efficient model adaptation across various ophthalmic tasks. Despite their success, current retinal foundation models are…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Kai Yu , Yang Zhou , Yang Bai , Zhi Da Soh , Xinxing Xu , Rick Siow Mong Goh , Ching-Yu Cheng , Yong Liu

Automated diagnosis based on color fundus photography is essential for large-scale glaucoma screening. However, existing deep learning models are typically data-driven and lack explicit integration of retinal anatomical knowledge, which…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Yuzhuo Zhou , Chi Liu , Sheng Shen , Zongyuan Ge , Fengshi Jing , Shiran Zhang , Yu Jiang , Anli Wang , Wenjian Liu , Feilong Yang , Tianqing Zhu , Xiaotong Han

The joint interpretation of multi-modal and multi-view fundus images is critical for retinopathy prevention, as different views can show the complete 3D eyeball field and different modalities can provide complementary lesion areas. Compared…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Yonghao Huang , Leiting Chen , Chuan Zhou

Our research is motivated by the urgent global issue of a large population affected by retinal diseases, which are evenly distributed but underserved by specialized medical expertise, particularly in non-urban areas. Our primary objective…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Deependra Singh , Saksham Agarwal , Subhankar Mishra

This study aimed to enhance disease classification accuracy from retinal fundus images by integrating fine-grained image features and global textual context using a novel multimodal deep learning architecture. Existing multimodal large…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Jason Jordan , Mohammadreza Akbari Lor , Peter Koulen , Mei-Ling Shyu , Shu-Ching Chen

Recent advancements in deep learning have shown significant potential for classifying retinal diseases using color fundus images. However, existing works predominantly rely exclusively on image data, lack interpretability in their…

图像与视频处理 · 电气工程与系统科学 2025-03-06 Deval Mehta , Yiwen Jiang , Catherine L Jan , Mingguang He , Kshitij Jadhav , Zongyuan Ge

Retinal blood vessel segmentation can extract clinically relevant information from fundus images. As manual tracing is cumbersome, algorithms based on Convolution Neural Networks have been developed. Such studies have used small publicly…

图像与视频处理 · 电气工程与系统科学 2024-06-24 Jeremiah Fadugba , Patrick Köhler , Lisa Koch , Petru Manescu , Philipp Berens

Over the past decade, generative models have achieved significant success in enhancement fundus images.However, the evaluation of these models still presents a considerable challenge. A comprehensive evaluation benchmark for fundus image…

图像与视频处理 · 电气工程与系统科学 2025-02-21 Wenhui Zhu , Xuanzhao Dong , Xin Li , Yujian Xiong , Xiwen Chen , Peijie Qiu , Vamsi Krishna Vasa , Zhangsihao Yang , Yi Su , Oana Dumitrascu , Yalin Wang

Foundation vision-language models are currently transforming computer vision, and are on the rise in medical imaging fueled by their very promising generalization capabilities. However, the initial attempts to transfer this new paradigm to…

计算机视觉与模式识别 · 计算机科学 2025-01-16 Julio Silva-Rodríguez , Hadi Chakor , Riadh Kobbi , Jose Dolz , Ismail Ben Ayed

Previous foundation models for fundus images were pre-trained with limited disease categories and knowledge base. Here we introduce a knowledge-rich vision-language model (RetiZero) that leverages knowledge from more than 400 fundus…

Existing multi-modal learning methods on fundus and OCT images mostly require both modalities to be available and strictly paired for training and testing, which appears less practical in clinical scenarios. To expand the scope of clinical…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Lehan Wang , Chongchong Qi , Chubin Ou , Lin An , Mei Jin , Xiangbin Kong , Xiaomeng Li

Recent advancements in pre-trained large foundation models (LFM) have yielded significant breakthroughs across various domains, including natural language processing and computer vision. These models have been particularly impactful in the…

图像与视频处理 · 电气工程与系统科学 2024-05-22 Ziqin Lin , Heng Li , Zinan Li , Huazhu Fu , Jiang Liu

This paper attacks an emerging challenge of multi-modal retinal disease recognition. Given a multi-modal case consisting of a color fundus photo (CFP) and an array of OCT B-scan images acquired during an eye examination, we aim to build a…

计算机视觉与模式识别 · 计算机科学 2021-09-28 Xirong Li , Yang Zhou , Jie Wang , Hailan Lin , Jianchun Zhao , Dayong Ding , Weihong Yu , Youxin Chen
‹ 上一页 1 2 3 10 下一页 ›