English
Related papers

Related papers: EyeFound: A Multimodal Generalist Foundation Model…

200 papers

Recent advances in generative AI have brought incredible breakthroughs in several areas, including medical imaging. These generative models have tremendous potential not only to help safely share medical data via synthetic datasets but also…

Medical AI assistants support doctors in disease diagnosis, medical image analysis, and report generation. However, they still face significant challenges in clinical use, including limited accuracy with multimodal content and insufficient…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Haonan Wang , Jiaji Mao , Lehan Wang , Qixiang Zhang , Marawan Elbatel , Yi Qin , Huijun Hu , Baoxun Li , Wenhui Deng , Weifeng Qin , Hongrui Li , Jialin Liang , Jun Shen , Xiaomeng Li

Multimodal Large Language Models (MLLMs) have tremendous potential to improve the accuracy, availability, and cost-effectiveness of healthcare by providing automated solutions or serving as aids to medical professionals. Despite promising…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Mohammad Shahab Sepehri , Zalan Fabian , Maryam Soltanolkotabi , Mahdi Soltanolkotabi

Artificial intelligence has shown the potential to improve diagnostic accuracy through medical image analysis for pneumonia diagnosis. However, traditional multimodal approaches often fail to address real-world challenges such as incomplete…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Jingyu Xu , Yang Wang

Recent advances in AI combine large language models (LLMs) with vision encoders that bring forward unprecedented technical capabilities to leverage for a wide range of healthcare applications. Focusing on the domain of radiology,…

Challenges in the field of retinal prostheses motivate the development of retinal models to accurately simulate Retinal Ganglion Cells (RGCs) responses. The goal of retinal prostheses is to enable blind individuals to solve complex,…

Image and Video Processing · Electrical Eng. & Systems 2022-02-08 Nikolas Papadopoulos , Nikos Melanitis , Antonio Lozano , Cristina Soto-Sanchez , Eduardo Fernandez , Konstantina S Nikita

Artificial intelligence (AI) has the potential to transform medical imaging by automating image analysis and accelerating clinical research. However, research and clinical use are limited by the wide variety of AI implementations and…

Existing multi-modal learning methods on fundus and OCT images mostly require both modalities to be available and strictly paired for training and testing, which appears less practical in clinical scenarios. To expand the scope of clinical…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Lehan Wang , Chongchong Qi , Chubin Ou , Lin An , Mei Jin , Xiangbin Kong , Xiaomeng Li

Multimodal large language models (MLLMs) have shown remarkable potential in various domains, yet their application in the medical field is hindered by several challenges. General-purpose MLLMs often lack the specialized knowledge required…

Artificial Intelligence · Computer Science 2025-09-29 Guanghao Zhu , Zhitian Hou , Zeyu Liu , Zhijie Sang , Congkai Xie , Hongxia Yang

Foundation models and vision-language pre-training have significantly advanced Vision-Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their application in domain-specific agricultural tasks,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Khang Nguyen Quoc , Phuong D. Dao , Luyl-Da Quach

Inability to express the confidence level and detect unseen classes has limited the clinical implementation of artificial intelligence in the real-world. We developed a foundation model with uncertainty estimation (FMUE) to detect 11…

Image and Video Processing · Electrical Eng. & Systems 2024-06-26 Yuanyuan Peng , Aidi Lin , Meng Wang , Tian Lin , Ke Zou , Yinglin Cheng , Tingkun Shi , Xulong Liao , Lixia Feng , Zhen Liang , Xinjian Chen , Huazhu Fu , Haoyu Chen

AI Foundation models are gaining traction in various applications, including medical fields like radiology. However, medical foundation models are often tested on limited tasks, leaving their generalisability and biases unexplored. We…

Federated learning (FL) has become a promising paradigm for collaborative medical image analysis, yet existing frameworks remain tightly coupled to task-specific backbones and are fragile under heterogeneous imaging modalities. Such…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Meilin Liu , Jiaying Wang , Jing Shan

Multimodal neuroimaging provides complementary insights for Alzheimer's disease diagnosis, yet clinical datasets frequently suffer from missing modalities. We propose ACADiff, a framework that synthesizes missing brain imaging modalities…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Rong Zhou , Houliang Zhou , Yao Su , Brian Y. Chen , Yu Zhang , Lifang He , Alzheimer's Disease Neuroimaging Initiative

Multi-modal foundation models are typically trained on millions of pairs of natural images and text captions, frequently obtained through web-crawling approaches. Although such models depict excellent generative capabilities, they do not…

Computer Vision and Pattern Recognition · Computer Science 2023-01-03 Pierre Chambon , Christian Bluethgen , Curtis P. Langlotz , Akshay Chaudhari

Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains,…

Foundation model, which is pre-trained on broad data and is able to adapt to a wide range of tasks, is advancing healthcare. It promotes the development of healthcare artificial intelligence (AI) models, breaking the contradiction between…

Computers and Society · Computer Science 2024-04-05 Yuting He , Fuxiang Huang , Xinrui Jiang , Yuxiang Nie , Minghao Wang , Jiguang Wang , Hao Chen

Foundation models for vision and language are the basis of AI applications across numerous sectors of society. The success of these models stems from their ability to mimic human capabilities, namely visual perception in vision models, and…

Human-Computer Interaction · Computer Science 2024-10-08 Matthew Berger , Shusen Liu

Artificial Intelligence (AI) has become commonplace to solve routine everyday tasks. Because of the exponential growth in medical imaging data volume and complexity, the workload on radiologists is steadily increasing. We project that the…

Multimodal learning has witnessed remarkable advancements in recent years, particularly with the integration of attention-based models, leading to significant performance gains across a variety of tasks. Parallel to this progress, the…

Machine Learning · Computer Science 2026-04-28 Md Raisul Kibria , Sébastien Lafond , Janan Arslan