English
Related papers

Related papers: MM-Retinal: Knowledge-Enhanced Foundational Pretra…

200 papers

Recent advances in multimodal foundation models have achieved state-of-the-art performance across a range of tasks. These breakthroughs are largely driven by new pre-training paradigms that leverage large-scale, unlabeled multimodal data,…

Machine Learning · Computer Science 2025-06-10 Xiaojun Shan , Qi Cao , Xing Han , Haofei Yu , Paul Pu Liang

Large language models such as BERT and the GPT series started a paradigm shift that calls for building general-purpose models via pre-training on large datasets, followed by fine-tuning on task-specific datasets. There is now a plethora of…

Computation and Language · Computer Science 2023-06-13 Jeremy Gwinnup , Kevin Duh

Retinal fundus photography is significant in diagnosing and monitoring retinal diseases. However, systemic imperfections and operator/patient-related factors can hinder the acquisition of high-quality retinal images. Previous efforts in…

Image and Video Processing · Electrical Eng. & Systems 2024-09-18 Xuanzhao Dong , Vamsi Krishna Vasa , Wenhui Zhu , Peijie Qiu , Xiwen Chen , Yi Su , Yujian Xiong , Zhangsihao Yang , Yanxi Chen , Yalin Wang

Fundus image quality is crucial for diagnosing eye diseases, but real-world conditions often result in blurred or unreadable images, increasing diagnostic uncertainty. To address these challenges, this study proposes RetinaRegen, a hybrid…

Image and Video Processing · Electrical Eng. & Systems 2025-02-28 Yuhan Tang , Yudian Wang , Weizhen Li , Ye Yue , Chengchang Pan , Honggang Qi

The mechanism of connecting multimodal signals through self-attention operation is a key factor in the success of multimodal Transformer networks in remote sensing data fusion tasks. However, traditional approaches assume access to all…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Yuxing Chen , Maofan Zhao , Lorenzo Bruzzone

Diabetic Retinopathy (DR) is a serious microvascular complication of diabetes, and one of the leading causes of vision loss worldwide. Although automated detection and grading, with Deep Learning (DL), can reduce the burden on…

Image and Video Processing · Electrical Eng. & Systems 2026-04-06 Shramana Dey , Zahir Khan , T. A. PramodKumar , B. Uma Shankar , Ashis K. Dhara , Ramachandran Rajalakshmi , Rajiv Raman , Sushmita Mitra

While image-text representation learning has become very popular in recent years, existing models tend to lack spatial awareness and have limited direct applicability for dense understanding tasks. For this reason, self-supervised…

Multimodal Large Language Models (MLLMs) have shown significant potential in medical image analysis. However, their capabilities in interpreting fundus images, a critical skill for ophthalmology, remain under-evaluated. Existing benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Qijie Wei , Kaiheng Qian , Xirong Li

Multimodal foundation models serve numerous applications at the intersection of vision and language. Still, despite being pretrained on extensive data, they become outdated over time. To keep models updated, research into continual…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Karsten Roth , Vishaal Udandarao , Sebastian Dziadzio , Ameya Prabhu , Mehdi Cherti , Oriol Vinyals , Olivier Hénaff , Samuel Albanie , Matthias Bethge , Zeynep Akata

Foundation models or pre-trained models have substantially improved the performance of various language, vision, and vision-language understanding tasks. However, existing foundation models can only perform the best in one type of tasks,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 Xinsong Zhang , Yan Zeng , Jipeng Zhang , Hang Li

Multimodal image fusion (MMIF) integrates information from different modalities to obtain a comprehensive image, aiding downstream tasks. However, existing research focuses on complementary information fusion and training strategies,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Dan He , Guofen Wang , Weisheng Li , Yucheng Shu , Wenbo Li , Lijian Yang , Yuping Huang , Feiyan Li

Foundational multimodal models pre-trained on large scale image-text pairs or video-text pairs or both have shown strong generalization abilities on downstream tasks. However unlike image-text models, pretraining video-text models is always…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Avinash Madasu , Anahita Bhiwandiwalla , Vasudev Lal

There is substantial interest in developing artificial intelligence systems to support radiologists across tasks ranging from segmentation to report generation. Existing computed tomography (CT) foundation models have largely focused on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Rubén Moreno-Aguado , Alba Magallón , Victor Moreno , Yingying Fang , Guang Yang

Despite the significant progress made by all-in-one models in universal image restoration, existing methods suffer from a generalization bottleneck in real-world scenarios, as they are mostly trained on small-scale synthetic datasets with…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Hao Li , Xiang Chen , Jiangxin Dong , Jinhui Tang , Jinshan Pan

Vision Foundation Models (VFMs) pretrained on massive datasets exhibit impressive performance on various downstream tasks, especially with limited labeled target data. However, due to their high inference compute cost, these models cannot…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Raviteja Vemulapalli , Hadi Pouransari , Fartash Faghri , Sachin Mehta , Mehrdad Farajtabar , Mohammad Rastegari , Oncel Tuzel

Diabetic retinopathy is the most important complication of diabetes. Early diagnosis of retinal lesions helps to avoid visual loss or blindness. Due to high-resolution and small-size lesion regions, applying existing methods, such as…

Computer Vision and Pattern Recognition · Computer Science 2019-01-21 Zizheng Yan , Xiaoguang Han , Changmiao Wang , Yuda Qiu , Zixiang Xiong , Shuguang Cui

Purpose: Convolutional neural networks can be trained to detect various conditions or patient traits based on retinal fundus photographs, some of which, such as the patient sex, are invisible to the expert human eye. Here we propose a…

Computer Vision and Pattern Recognition · Computer Science 2023-01-18 Parsa Delavari , Gulcenur Ozturan , Ozgur Yilmaz , Ipek Oruc

Retinopathy represents a group of retinal diseases that, if not treated timely, can cause severe visual impairments or even blindness. Many researchers have developed autonomous systems to recognize retinopathy via fundus and optical…

Image and Video Processing · Electrical Eng. & Systems 2021-11-05 Taimur Hassan , Bilal Hassan , Muhammad Usman Akram , Shahrukh Hashmi , Abdel Hakim Taguri , Naoufel Werghi

Artificial Intelligence in medicine is traditionally limited by the lack of massive training datasets. Foundation models, pre-trained models that can be adapted to downstream tasks with small datasets, could alleviate this problem.…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Justin Engelmann , Miguel O. Bernabeu

As real-world knowledge continues to evolve, the parametric knowledge acquired by multimodal models during pretraining becomes increasingly difficult to remain consistent with real-world knowledge. Existing research on multimodal knowledge…

Computation and Language · Computer Science 2026-03-17 Baochen Fu , Yuntao Du , Cheng Chang , Baihao Jin , Wenzhi Deng , Muhao Xu , Hongmei Yan , Weiye Song , Yi Wan
‹ Prev 1 4 5 6 7 8 10 Next ›