English
Related papers

Related papers: Evaluating Vision Language Models (VLMs) for Radio…

200 papers

CLIP outperforms self-supervised models like DINO as vision encoders for vision-language models (VLMs), but it remains unclear whether this advantage stems from CLIP's language supervision or its much larger training data. To disentangle…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Yiming Liu , Yuhui Zhang , Dhruba Ghosh , Ludwig Schmidt , Serena Yeung-Levy

Vision-language pre-training has recently gained popularity as it allows learning rich feature representations using large-scale data sources. This paradigm has quickly made its way into the medical image analysis community. In particular,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Julio Silva-Rodríguez , Jose Dolz , Ismail Ben Ayed

Clinicians spend a significant amount of time reviewing medical images and transcribing their findings regarding patient diagnosis, referral and treatment in text form. Vision-language models (VLMs), which automatically interpret images and…

Accurate segmentation of pulmonary structures iscrucial in clinical diagnosis, disease study, and treatment planning. Significant progress has been made in deep learning-based segmentation techniques, but most require much labeled data for…

Computer Vision and Pattern Recognition · Computer Science 2025-01-08 Xiaotong Guo , Deqian Yang , Dan Wang , Haochen Zhao , Yuan Li , Zhilin Sui , Tao Zhou , Lijun Zhang , Yanda Meng

In this paper, we introduce CheXOFA, a new pre-trained vision-language model (VLM) for the chest X-ray domain. Our model is initially pre-trained on various multimodal datasets within the general domain before being transferred to the chest…

Computation and Language · Computer Science 2023-07-17 Gangwoo Kim , Hajung Kim , Lei Ji , Seongsu Bae , Chanhwi Kim , Mujeen Sung , Hyunjae Kim , Kun Yan , Eric Chang , Jaewoo Kang

Automated radiology report generation is key for reducing radiologist workload and improving diagnostic consistency, yet generating accurate reports for 3D medical imaging remains challenging. Existing vision-language models face two…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Pengcheng Shi , Minghui Zhang , Kehan Song , Jiaqi Liu , Yun Gu , Xinglin Zhang

Modern deep learning offers powerful tools for automated retinal screening, but it remains unclear how different visual model families compare in realistic multi-disease settings and under domain shift. In this work, we benchmark twelve…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Durjoy Dey , Aymane Ajbar , Yuhong Yan

Vision-language foundation models (VLMs) have shown impressive performance in guiding image generation through text, with emerging applications in medical imaging. In this work, we are the first to investigate the question: 'Can fine-tuned…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Amar Kumar , Anita Kriz , Barak Pertzov , Tal Arbel

Lung cancer remains one of the leading causes of cancer-related mortality worldwide. A crucial challenge for early diagnosis is differentiating uncertain cases with similar visual characteristics and closely annotation scores. In clinical…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Sadaf Khademi , Mehran Shabanpour , Reza Taleei , Anastasia Oikonomou , Arash Mohammadi

The deep learning field is converging towards the use of general foundation models that can be easily adapted for diverse tasks. While this paradigm shift has become common practice within the field of natural language processing, progress…

Computer Vision and Pattern Recognition · Computer Science 2023-11-15 Joana Palés Huix , Adithya Raju Ganeshan , Johan Fredin Haslum , Magnus Söderberg , Christos Matsoukas , Kevin Smith

Large Language Models (LLMs), known for their versatility in textual data, are increasingly being explored for their potential to enhance medical image segmentation, a crucial task for accurate diagnostic imaging. This study explores…

Image and Video Processing · Electrical Eng. & Systems 2025-08-20 Gurucharan Marthi Krishna Kumar , Aman Chadha , Janine Mendola , Amir Shmuel

The clinical adoption of artificial intelligence (AI) in medical imaging requires models that are both diagnostically accurate and interpretable to clinicians. While current multimodal biomedical foundation models prioritize performance,…

Vision-language models (VLMs) have shown considerable potential in digital pathology, yet their effectiveness remains limited for fine-grained, disease-specific classification tasks such as distinguishing between glomerular subtypes. The…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Zhenhao Guo , Rachit Saluja , Tianyuan Yao , Quan Liu , Yuankai Huo , Benjamin Liechty , David J. Pisapia , Kenji Ikemura , Mert R. Sabuncu , Yihe Yang , Ruining Deng

Deep learning methods have demonstrated promising results in predicting BI-RADS scores from mammography images. However, the interpretation of these images can vary, leading to discrepancies even among radiologists. Given the inherent…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Halil Ibrahim Gulluk , Olivier Gevaert

Artificial intelligence (AI) has significant potential in healthcare applications, but its training and deployment faces challenges due to healthcare's diverse data, complex tasks, and the need to preserve privacy. Foundation models that…

Most natural language tasks in the radiology domain use language models pre-trained on biomedical corpus. There are few pretrained language models trained specifically for radiology, and fewer still that have been trained in a low data…

Computation and Language · Computer Science 2023-06-06 Rikhiya Ghosh , Sanjeev Kumar Karn , Manuela Daniela Danu , Larisa Micu , Ramya Vunikili , Oladimeji Farri

Recently large vision-language models have shown potential when interpreting complex images and generating natural language descriptions using advanced reasoning. Medicine's inherently multimodal nature incorporating scans and text-based…

Image and Video Processing · Electrical Eng. & Systems 2024-07-15 Naman Sharma

In this study, we aim to initiate the development of Radiology Foundation Model, termed as RadFM. We consider the construction of foundational models from three perspectives, namely, dataset construction, model design, and thorough…

Computer Vision and Pattern Recognition · Computer Science 2023-11-17 Chaoyi Wu , Xiaoman Zhang , Ya Zhang , Yanfeng Wang , Weidi Xie

Automated chest radiographs interpretation requires both accurate disease classification and detailed radiology report generation, presenting a significant challenge in the clinical workflow. Current approaches either focus on…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Difei Gu , Yunhe Gao , Yang Zhou , Mu Zhou , Dimitris Metaxas

Vision foundation models (VFMs) are pre-trained on extensive image datasets to learn general representations for diverse types of data. These models can subsequently be fine-tuned for specific downstream tasks, significantly boosting…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Shansong Wang , Mojtaba Safari , Qiang Li , Chih-Wei Chang , Richard LJ Qiu , Justin Roper , David S. Yu , Xiaofeng Yang