English
Related papers

Related papers: Generalist Foundation Models from a Multimodal Dat…

200 papers

Understanding radiologists' eye movement during Computed Tomography (CT) reading is crucial for developing effective interpretable computer-aided diagnosis systems. However, CT research in this area has been limited by the lack of publicly…

Objective: While recent advances in text-conditioned generative models have enabled the synthesis of realistic medical images, progress has been largely confined to 2D modalities such as chest X-rays. Extending text-to-image generation to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Daniele Molino , Camillo Maria Caruso , Filippo Ruffini , Paolo Soda , Valerio Guarrasi

Electrocardiography (ECG) is central to cardiovascular care, but conventional AI models are often restricted to common arrhythmias and may generalize poorly across populations or clinically subtle diseases. We developed ECG Contrastive…

The latest breakthroughs in large vision-language models, such as Bard and GPT-4, have showcased extraordinary abilities in performing a wide range of tasks. Such models are trained on massive datasets comprising billions of public…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Omkar Thawakar , Abdelrahman Shaker , Sahal Shaji Mullappilly , Hisham Cholakkal , Rao Muhammad Anwer , Salman Khan , Jorma Laaksonen , Fahad Shahbaz Khan

Ultrasound foundation models have achieved strong performance on structured prediction tasks but remain exclusively vision-based, limiting zero-shot and few-shot transfer to novel tasks where task-specific annotation is scarce. We address…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Zhuoyang Lyu , Yiyang Zhang , Tongxin Wang , Ruirui Lan

Contrastive Language-Image Pre-training (CLIP) demonstrates strong potential in medical image analysis but requires substantial data and computational resources. Due to these restrictions, existing CLIP applications in medical imaging focus…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Yuexi Du , John Onofrey , Nicha C. Dvornek

Diversity in data is critical for the successful training of deep learning models. Leveraged by a recurrent generative adversarial network, we propose the CT-SGAN model that generates large-scale 3D synthetic CT-scan volumes ($\geq…

Image and Video Processing · Electrical Eng. & Systems 2021-11-08 Ahmad Pesaranghader , Yiping Wang , Mohammad Havaei

The clinical adoption of artificial intelligence (AI) in medical imaging requires models that are both diagnostically accurate and interpretable to clinicians. While current multimodal biomedical foundation models prioritize performance,…

Cancers identified in CT scans are usually accompanied by detailed radiology reports, but publicly available CT datasets often lack these essential reports. This absence limits their usefulness for developing accurate report generation AI.…

Image and Video Processing · Electrical Eng. & Systems 2025-08-20 Pedro R. A. S. Bassi , Mehmet Can Yavuz , Kang Wang , Xiaoxi Chen , Wenxuan Li , Sergio Decherchi , Andrea Cavalli , Yang Yang , Alan Yuille , Zongwei Zhou

Vision-Language Models (VLMs) trained via contrastive learning have achieved notable success in natural image tasks. However, their application in the medical domain remains limited due to the scarcity of openly accessible, large-scale…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Muhammad Uzair Khattak , Shahina Kunhimon , Muzammal Naseer , Salman Khan , Fahad Shahbaz Khan

We developed a rich dataset of Chest X-Ray (CXR) images to assist investigators in artificial intelligence. The data were collected using an eye tracking system while a radiologist reviewed and reported on 1,083 CXR images. The dataset…

Since the release of the original CheXpert paper five years ago, CheXpert has become one of the most widely used and cited clinical AI datasets. The emergence of vision language models has sparked an increase in demands for sharing reports…

Recent advancements in artificial intelligence (AI) have precipitated significant breakthroughs in healthcare, particularly in refining diagnostic procedures. However, previous studies have often been constrained to limited functionalities.…

Artificial Intelligence · Computer Science 2024-07-08 Asma Alkhaldi , Raneem Alnajim , Layan Alabdullatef , Rawan Alyahya , Jun Chen , Deyao Zhu , Ahmed Alsinan , Mohamed Elhoseiny

Echocardiography is the most widely used cardiac imaging modality, capturing ultrasound video data to assess cardiac structure and function. Artificial intelligence (AI) in echocardiography has the potential to streamline manual tasks and…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Milos Vukadinovic , Xiu Tang , Neal Yuan , Paul Cheng , Debiao Li , Susan Cheng , Bryan He , David Ouyang

As advances in large language models (LLMs) and multimodal techniques continue to mature, the development of general-purpose multimodal large language models (MLLMs) has surged, offering significant applications in interpreting natural…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Yuxuan Sun , Chenglu Zhu , Sunyi Zheng , Kai Zhang , Lin Sun , Zhongyi Shui , Yunlong Zhang , Honglin Li , Lin Yang

The deployment of artificial intelligence in medical imaging is hindered by high computational complexity and resource-intensive processing of volumetric data. Although chest computed tomography (CT) volumes offer richer diagnostic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Shadid Yousuf , S. M. Mahbubur Rahman , Mohammed Imamul Hassan Bhuiyan

Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains,…

GenerateCT, the first approach to generating 3D medical imaging conditioned on free-form medical text prompts, incorporates a text encoder and three key components: a novel causal vision transformer for encoding 3D CT volumes, a text-image…

Ultrasound imaging is widely used in clinical diagnostics due to its real-time capability and radiation-free nature. However, existing vision-language pre-training models, such as CLIP, are primarily designed for other modalities, and are…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Jiayun Jin , Haolong Chai , Xueying Huang , Xiaoqing Guo , Zengwei Zheng , Zhan Zhou , Junmei Wang , Xinyu Wang , Jie Liu , Binbin Zhou