English
Related papers

Related papers: Generalist Foundation Models from a Multimodal Dat…

200 papers

Medical imaging is essential in modern radiotherapy, supporting diagnosis, treatment planning, and monitoring. Synthetic imaging, particularly synthetic computed tomography (sCT), is gaining traction in radiotherapy. The SynthRAD2025…

ChatGPT is attracting a cross-field interest as it provides a language interface with remarkable conversational competency and reasoning capabilities across many domains. However, since ChatGPT is trained with languages, it is currently not…

Computer Vision and Pattern Recognition · Computer Science 2023-03-09 Chenfei Wu , Shengming Yin , Weizhen Qi , Xiaodong Wang , Zecheng Tang , Nan Duan

Generative models have revolutionized Artificial Intelligence (AI), particularly in multimodal applications. However, adapting these models to the medical domain poses unique challenges due to the complexity of medical data and the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Daniele Molino , Francesco di Feola , Linlin Shen , Paolo Soda , Valerio Guarrasi

The widespread use of chest X-rays (CXRs), coupled with a shortage of radiologists, has driven growing interest in automated CXR analysis and AI-assisted reporting. While existing vision-language models (VLMs) show promise in specific tasks…

Large language models, notably utilizing Transformer architectures, have emerged as powerful tools due to their scalability and ability to process large amounts of data. Dosovitskiy et al. expanded this architecture to introduce Vision…

Image and Video Processing · Electrical Eng. & Systems 2024-06-04 Ananya Jain , Aviral Bhardwaj , Kaushik Murali , Isha Surani

AI in dermatology is evolving at a rapid pace but the major limitation to training trustworthy classifiers is the scarcity of data with ground-truth concept level labels, which are meta-labels semantically meaningful to humans. Foundation…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Soham Gadgil , Mahtab Bigverdi

The tremendous success of CLIP (Radford et al., 2021) has promoted the research and application of contrastive learning for vision-language pretraining. In this work, we construct a large-scale dataset of image-text pairs in Chinese, where…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 An Yang , Junshu Pan , Junyang Lin , Rui Men , Yichang Zhang , Jingren Zhou , Chang Zhou

Contrastive Language-Image Pre-training (CLIP), a simple yet effective pre-training paradigm, successfully introduces text supervision to vision models. It has shown promising results across various tasks due to its generalizability and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Zihao Zhao , Yuxiao Liu , Han Wu , Mei Wang , Yonghao Li , Sheng Wang , Lin Teng , Disheng Liu , Zhiming Cui , Qian Wang , Dinggang Shen

The chest X-ray is one of the most commonly accessible radiological examinations for screening and diagnosis of many lung diseases. A tremendous number of X-ray imaging studies accompanied by radiological reports are accumulated and stored…

Computer Vision and Pattern Recognition · Computer Science 2019-02-01 Xiaosong Wang , Yifan Peng , Le Lu , Zhiyong Lu , Mohammadhadi Bagheri , Ronald M. Summers

Multi-modal models require aligned, shared embedding spaces. However, common CLIP-based approaches need large amounts of samples and do not natively support 3D or tabular data, both of which are crucial in the medical domain. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Jakob Krogh Petersen , Valdemar Licht , Mads Nielsen , Asbjørn Munk

Multimodal deep learning foundation models can learn the relationship between images and text. In the context of medical imaging, mapping images to language concepts reflects the clinical task of diagnostic image interpretation, however…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Matthew Christensen , Milos Vukadinovic , Neal Yuan , David Ouyang

The outbreak of COVID-19 disease caused more than 100,000 deaths so far in the USA alone. It is necessary to conduct an initial screening of patients with the symptoms of COVID-19 disease to control the spread of the disease. However, it is…

Image and Video Processing · Electrical Eng. & Systems 2020-10-30 Md Manjurul Ahsan , Kishor Datta Gupta , Mohammad Maminur Islam , Sajib Sen , Md. Lutfar Rahman , Mohammad Shakhawat Hossain

We present brat (brain report alignment transformer), a multi-view representation learning framework for brain magnetic resonance imaging (MRI) trained on MRIs paired with clinical reports. Brain MRIs present unique challenges due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Maxime Kayser , Maksim Gridnev , Wanting Wang , Max Bain , Aneesh Rangnekar , Avijit Chatterjee , Aleksandr Petrov , Harini Veeraraghavan , Nathaniel C. Swinburne

Radiologic diagnostic errors-under-reading errors, inattentional blindness, and communication failures-remain prevalent in clinical practice. These issues often stem from missed localized abnormalities, limited global context, and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Yuheng Li , Yenho Chen , Yuxiang Lai , Jike Zhong , Vanessa Wildman , Xiaofeng Yang

After more than two years since the beginning of the COVID-19 pandemic, the pressure of this crisis continues to devastate globally. The use of chest X-ray (CXR) imaging as a complementary screening strategy to RT-PCR testing is not only…

Image and Video Processing · Electrical Eng. & Systems 2022-11-22 Maya Pavlova , Tia Tuinstra , Hossein Aboutalebi , Andy Zhao , Hayden Gunraj , Alexander Wong

Coronary angiography is the reference standard for evaluating coronary artery disease, yet visual interpretation remains variable between readers. Existing artificial intelligence methods typically analyze single frames or projections and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Sarra Harrabi , Yichen Wu , Geoffrey H. Tison , Minhaj Ansari , Milos Vukadinovic , David Ouyang , Joshua P. Barrios , Jacques Delfrate , Robert Avram

Purpose: to optimize a pipeline of clinical data gathering and CT images processing implemented during the COVID-19 pandemic crisis and to develop artificial intelligence model for different of viral pneumonia. Methods: 1028 chest CT image…

Medical Physics · Physics 2021-09-30 Giulia Zorzi , Luca Berta , Stefano Carrazza , Alberto Torresin

Contrastive Language-Image Pre-training (CLIP) on large-scale image-caption datasets learns representations that can achieve remarkable zero-shot generalization. However, such models require a massive amount of pre-training data. Improving…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Siddharth Joshi , Arnav Jain , Ali Payani , Baharan Mirzasoleiman

Pathology images are crucial for diagnosing and managing various diseases by visualizing cellular and tissue-level abnormalities. Recent advancements in artificial intelligence (AI), particularly multimodal models like ChatGPT, have shown…

Human-Computer Interaction · Computer Science 2024-09-25 Mianxin Liu , Jianfeng Wu , Fang Yan , Hongjun Li , Wei Wang , Shaoting Zhang , Zhe Wang

Detecting hateful content in multimodal memes presents unique challenges, as harmful messages often emerge from the complex interplay between benign images and text. We propose GatedCLIP, a Vision-Language model that enhances CLIP's…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Yingying Guo , Ke Zhang , Zirong Zeng