English
Related papers

Related papers: MIRAGE: Multimodal foundation model and benchmark …

200 papers

Open-vocabulary semantic segmentation (OVS) aims to segment images of arbitrary categories specified by class labels or captions. However, most previous best-performing methods, whether pixel grouping methods or region recognition methods,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Yuan Wang , Rui Sun , Naisong Luo , Yuwen Pan , Tianzhu Zhang

Radiology plays an integral role in modern medicine, yet rising imaging volumes have far outpaced workforce growth. Foundation models offer a path toward assisting with the full spectrum of radiology tasks, but existing medical models…

Objective Structured Clinical Examinations (OSCEs) are widely used to assess medical students' communication skills, but scoring interview-based assessments is time-consuming and potentially subject to human bias. This study explored the…

Computation and Language · Computer Science 2025-05-16 Jadon Geathers , Yann Hicke , Colleen Chan , Niroop Rajashekar , Justin Sewell , Susannah Cornes , Rene F. Kizilcec , Dennis Shung

Pretraining on large-scale, in-domain datasets grants histopathology foundation models (FM) the ability to learn task-agnostic data representations, enhancing transfer learning on downstream tasks. In computational pathology, automated…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Pablo Meseguer , Rocío del Amor , Valery Naranjo

The advent of foundation models signals a new era in artificial intelligence. The Segment Anything Model (SAM) is the first foundation model for image segmentation. In this study, we evaluate SAM's ability to segment features from eye…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Virmarie Maquiling , Sean Anthony Byrne , Diederick C. Niehorster , Marcus Nyström , Enkelejda Kasneci

Diabetic retinopathy (DR) is a leading cause of vision loss, requiring early and accurate assessment to prevent irreversible damage. Spectral Domain Optical Coherence Tomography (SD-OCT) enables high-resolution retinal imaging, but…

Image and Video Processing · Electrical Eng. & Systems 2025-07-15 S. Chen , D. Ma , M. Raviselvan , S. Sundaramoorthy , K. Popuri , M. J. Ju , M. V. Sarunic , D. Ratra , M. F. Beg

Optical coherence tomography (OCT) is a non-invasive, micrometer-scale imaging modality that has become a clinical standard in ophthalmology. By raster-scanning the retina, sequential cross-sectional image slices are acquired to generate…

Image and Video Processing · Electrical Eng. & Systems 2024-02-20 Stefan Ploner , Jungeun Won , Julia Schottenhamml , Jessica Girgis , Kenneth Lam , Nadia Waheed , James Fujimoto , Andreas Maier

Segmentation is vital for ophthalmology image analysis. But its various modal images hinder most of the existing segmentation algorithms applications, as they rely on training based on a large number of labels or hold weak generalization…

Computer Vision and Pattern Recognition · Computer Science 2023-04-27 Zhongxi Qiu , Yan Hu , Heng Li , Jiang Liu

Optical Coherence Tomography (OCT) provides a unique ability to image the eye retina in 3D at micrometer resolution and gives ophthalmologist the ability to visualize retinal diseases such as Age-Related Macular Degeneration (AMD). While…

Computer Vision and Pattern Recognition · Computer Science 2016-10-13 Stefanos Apostolopoulos , Carlos Ciller , Sandro I. De Zanet , Sebastian Wolf , Raphael Sznitman

In the latest advancements in multimodal learning, effectively addressing the spatial and semantic losses of visual data after encoding remains a critical challenge. This is because the performance of large multimodal models is positively…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Shaojun E , Yuchen Yang , Jiaheng Wu , Yan Zhang , Tiejun Zhao , Ziyan Chen

Medical foundation models have shown promise in controlled benchmarks, yet widespread deployment remains hindered by reliance on task-specific fine-tuning. Here, we introduce DermFM-Zero, a dermatology vision-language foundation model…

The manufacturing sector is increasingly adopting Multimodal Large Language Models (MLLMs) to transition from simple perception to autonomous execution, yet current evaluations fail to reflect the rigorous demands of real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Xiangru Jian , Hao Xu , Wei Pang , Xinjian Zhao , Chengyu Tao , Qixin Zhang , Xikun Zhang , Chao Zhang , Guanzhi Deng , Alex Xue , Juan Du , Tianshu Yu , Garth Tarr , Linqi Song , Qiuzhuang Sun , Dacheng Tao

Optical coherence tomography (OCT) and scanning laser ophthalmoscopy (SLO) of the eye has become essential to ophthalmology and the emerging field of oculomics, thus requiring a need for transparent, reproducible, and rapid analysis of this…

Glaucoma causes irreversible vision loss due to damage to the optic nerve, and there is no cure for glaucoma.OCT imaging modality is an essential technique for assessing glaucomatous damage since it aids in quantifying fundus structures. To…

Computer Vision and Pattern Recognition · Computer Science 2022-08-01 Huihui Fang , Fei Li , Huazhu Fu , Junde Wu , Xiulan Zhang , Yanwu Xu

Artificial intelligence (AI) tools for radiology are commonly unmonitored once deployed. The lack of real-time case-by-case assessments of AI prediction confidence requires users to independently distinguish between trustworthy and…

Optical Coherence Tomography (OCT) is the primary imaging modality for detecting pathological biomarkers associated to retinal diseases such as Age-Related Macular Degeneration. In practice, clinical diagnosis and treatment strategies are…

Computer Vision and Pattern Recognition · Computer Science 2019-07-17 Thomas Kurmann , Pablo Márquez-Neila , Siqing Yu , Marion Munk , Sebastian Wolf , Raphael Sznitman

Multimodal Magnetic Resonance Imaging (MRI) provides essential complementary information for analyzing brain tumor subregions. While methods using four common MRI modalities for automatic segmentation have shown success, they often face…

Image and Video Processing · Electrical Eng. & Systems 2024-11-14 Runze Cheng , Zhongao Sun , Ye Zhang , Chun Li

Multimodal large language models (MLLMs) show promise in tasks like visual question answering (VQA) but still face challenges in multimodal reasoning. Recent works adapt agentic frameworks or chain-of-thought (CoT) reasoning to improve…

Artificial Intelligence · Computer Science 2025-03-12 Zhuo Zhi , Chen Feng , Adam Daneshmend , Mine Orlu , Andreas Demosthenous , Lu Yin , Da Li , Ziquan Liu , Miguel R. D. Rodrigues

Large Language Models (LLMs) have shown remarkable capabilities in environmental perception, reasoning-based decision-making, and simulating complex human behaviors, particularly in interactive role-playing contexts. This paper introduces…

Computation and Language · Computer Science 2026-01-21 Yin Cai , Zhouhong Gu , Zhaohan Du , Zheyu Ye , Shaosheng Cao , Yiqian Xu , Hongwei Feng , Ping Chen
‹ Prev 1 8 9 10 Next ›