English
Related papers

Related papers: A Survey of Multimodal Ophthalmic Diagnostics: Fro…

200 papers

Vision-threatening eye diseases pose a major global health burden, with timely diagnosis limited by workforce shortages and restricted access to specialized care. While multimodal large language models (MLLMs) show promise for medical image…

Artificial intelligence (AI) is vital in ophthalmology, tackling tasks like diagnosis, classification, and visual question answering (VQA). However, existing AI models in this domain often require extensive annotation and are task-specific,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Danli Shi , Weiyi Zhang , Xiaolan Chen , Yexin Liu , Jiancheng Yang , Siyu Huang , Yih Chung Tham , Yingfeng Zheng , Mingguang He

The prevalence of vision-threatening eye diseases is a significant global burden, with many cases remaining undiagnosed or diagnosed too late for effective treatment. Large vision-language models (LVLMs) have the potential to assist in…

Computer Vision and Pattern Recognition · Computer Science 2025-02-06 Zhenyue Qin , Yu Yin , Dylan Campbell , Xuansheng Wu , Ke Zou , Yih-Chung Tham , Ninghao Liu , Xiuzhen Zhang , Qingyu Chen

Foundation models, large-scale, pre-trained deep-learning models adapted to a wide range of downstream tasks have gained significant interest lately in various deep-learning problems undergoing a paradigm shift with the rise of these…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Bobby Azad , Reza Azad , Sania Eskandari , Afshin Bozorgpour , Amirhossein Kazerouni , Islem Rekik , Dorit Merhof

This paper presents a comprehensive survey of the taxonomy and evolution of multimodal foundation models that demonstrate vision and vision-language capabilities, focusing on the transition from specialist models to general-purpose…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Chunyuan Li , Zhe Gan , Zhengyuan Yang , Jianwei Yang , Linjie Li , Lijuan Wang , Jianfeng Gao

Deep Learning has implemented a wide range of applications and has become increasingly popular in recent years. The goal of multimodal deep learning (MMDL) is to create models that can process and link information using various modalities.…

Machine Learning · Computer Science 2022-02-21 Jabeen Summaira , Xi Li , Amin Muhammad Shoib , Jabbar Abdul

Deep Learning has implemented a wide range of applications and has become increasingly popular in recent years. The goal of multimodal deep learning is to create models that can process and link information using various modalities. Despite…

Computer Vision and Pattern Recognition · Computer Science 2021-05-25 Jabeen Summaira , Xi Li , Amin Muhammad Shoib , Songyuan Li , Jabbar Abdul

Retinal diseases spanning a broad spectrum can be effectively identified and diagnosed using complementary signals from multimodal data. However, multimodal diagnosis in ophthalmic practice is typically challenged in terms of data…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Lu Zhang , Huizhen Yu , Zuowei Wang , Fu Gui , Yatu Guo , Wei Zhang , Mengyu Jia

The rapid advancement of remote sensing foundation models, particularly vision and multimodal models, has significantly enhanced the capabilities of intelligent geospatial data interpretation. These models combine various data modalities,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Ziyue Huang , Hongxi Yan , Qiqi Zhan , Shuai Yang , Mingming Zhang , Chenkai Zhang , YiMing Lei , Zeming Liu , Qingjie Liu , Yunhong Wang

Recent advancements in deep learning have significantly revolutionized the field of clinical diagnosis and treatment, offering novel approaches to improve diagnostic precision and treatment efficacy across diverse clinical domains, thus…

Artificial Intelligence · Computer Science 2024-12-04 Kai Sun , Siyan Xue , Fuchun Sun , Haoran Sun , Yu Luo , Ling Wang , Siyuan Wang , Na Guo , Lei Liu , Tian Zhao , Xinzhou Wang , Lei Yang , Shuo Jin , Jun Yan , Jiahong Dong

Computer-assisted diagnostic and prognostic systems of the future should be capable of simultaneously processing multimodal data. Multimodal deep learning (MDL), which involves the integration of multiple sources of data, such as images and…

Computer Vision and Pattern Recognition · Computer Science 2023-10-20 Zhaoyi Sun , Mingquan Lin , Qingqing Zhu , Qianqian Xie , Fei Wang , Zhiyong Lu , Yifan Peng

Foundation models have emerged as a powerful paradigm in computational pathology (CPath), enabling scalable and generalizable analysis of histopathological images. While early developments centered on uni-modal models trained solely on…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Dong Li , Guihong Wan , Xintao Wu , Xinyu Wu , Xiaohui Chen , Yi He , Christine G. Lian , Peter K. Sorger , Yevgeniy R. Semenov , Chen Zhao

This survey provides a comprehensive overview of recent advances in multimodal alignment and fusion within the field of machine learning, driven by the increasing availability and diversity of data modalities such as text, images, audio,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Songtao Li , Hao Tang

Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities with reinforcement learning paradigm. Although several multimodal reasoning models have been explored in the medical domain, most of them…

Artificial Intelligence · Computer Science 2025-09-11 Ruiqi Wu , Yuang Yao , Tengfei Ma , Chenran Zhang , Na Su , Tao Zhou , Geng Chen , Wen Fan , Yi Zhou

Vision impairment affects millions globally, and early detection is critical to preventing irreversible vision loss. Ophthalmology workflows require clinicians to integrate medical images, structured clinical data, and free-text notes to…

The advent of foundation models has heralded a new era in medical artificial intelligence (AI), enabling the extraction of generalizable representations from large-scale unlabeled datasets. However, current ophthalmic AI paradigms are…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Tienyu Chang , Zhen Chen , Renjie Liang , Jinyu Ding , Jie Xu , Sunu Mathew , Amir Reza Hajrasouliha , Andrew J. Saykin , Ruogu Fang , Yu Huang , Jiang Bian , Qingyu Chen

Large multimodal language models (LMMs) have achieved significant success in general domains. However, due to the significant differences between medical images and text and general web content, the performance of LMMs in medical scenarios…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Weihao Gao , Zhuo Deng , Zhiyuan Niu , Fuju Rong , Chucheng Chen , Zheng Gong , Wenze Zhang , Daimin Xiao , Fang Li , Zhenjie Cao , Zhaoyi Ma , Wenbin Wei , Lan Ma

Language models have recently advanced into the realm of reasoning, yet it is through multimodal reasoning that we can fully unlock the potential to achieve more comprehensive, human-like cognitive capabilities. This survey provides a…

Computation and Language · Computer Science 2025-03-25 Zhiyu Lin , Yifei Gao , Xian Zhao , Yunfan Yang , Jitao Sang

Multimodal data modeling has emerged as a powerful approach in clinical research, enabling the integration of diverse data types such as imaging, genomics, wearable sensors, and electronic health records. Despite its potential to improve…

With advanced imaging, sequencing, and profiling technologies, multiple omics data become increasingly available and hold promises for many healthcare applications such as cancer diagnosis and treatment. Multimodal learning for integrative…

Genomics · Quantitative Biology 2022-12-20 Sina Tabakhi , Mohammod Naimul Islam Suvon , Pegah Ahadian , Haiping Lu
‹ Prev 1 2 3 10 Next ›