English
Related papers

Related papers: CoCa-CXR: Contrastive Captioners Learn Strong Temp…

200 papers

Medical image interpretation is central to most clinical applications such as disease diagnosis, treatment planning, and prognostication. In clinical practice, radiologists examine medical images and manually compile their findings into…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Nurbanu Aksoy , Nishant Ravikumar , Alejandro F Frangi

In this work, we present an approach, which we call Embeddings for Language/Image-aligned X-Rays, or ELIXR, that leverages a language-aligned image encoder combined or grafted onto a fixed LLM, PaLM 2, to perform a broad range of chest…

A radiology report comprises several sections, including the Findings and Impression of the diagnosis. Automatically generating the Impression from the Findings is crucial for reducing radiologists' workload and improving diagnostic…

Recent research demonstrates that deep learning models are capable of precisely extracting bio-information (e.g. race, gender and age) from patients' Chest X-Rays (CXRs). In this paper, we further show that deep learning models are also…

Image and Video Processing · Electrical Eng. & Systems 2023-05-02 Hao Liang , Kevin Ni , Guha Balakrishnan

Recently, multi-modal vision-language foundation models have gained significant attention in the medical field. While these models offer great opportunities, they still face crucial challenges, such as the requirement for fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Weijian Huang , Cheng Li , Hong-Yu Zhou , Hao Yang , Jiarun Liu , Yong Liang , Hairong Zheng , Shaoting Zhang , Shanshan Wang

Automatic radiology reporting has great clinical potential to relieve radiologists from heavy workloads and improve diagnosis interpretation. Recently, researchers have enhanced data-driven neural networks with medical knowledge graphs to…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Mingjie Li , Bingqian Lin , Zicong Chen , Haokun Lin , Xiaodan Liang , Xiaojun Chang

Medical imaging analysis plays a critical role in the diagnosis and treatment of various medical conditions. This paper focuses on chest X-ray images and their corresponding radiological reports. It presents a new model that learns a joint…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Gefen Dawidowicz , Elad Hirsch , Ayellet Tal

The image captioning task is increasingly prevalent in artificial intelligence applications for medicine. One important application is clinical report generation from chest radiographs. The clinical writing of unstructured reports is time…

Image and Video Processing · Electrical Eng. & Systems 2022-05-09 Edward Vendrow , Ethan Schonfeld

Accurate and rapid detection of COVID-19 pneumonia is crucial for optimal patient treatment. Chest X-Ray (CXR) is the first line imaging test for COVID-19 pneumonia diagnosis as it is fast, cheap and easily accessible. Inspired by the…

Image and Video Processing · Electrical Eng. & Systems 2023-02-20 Xin Zhang , Liangxiu Han , Tam Sobeih , Lianghao Han , Nina Dempsey , Symeon Lechareas , Ascanio Tridente , Haoming Chen , Stephen White

The COVID-19 pandemic has strained global public health, necessitating accurate diagnosis and intervention to control disease spread and reduce mortality rates. This paper introduces an interpretable deep survival prediction model designed…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Zhusi Zhong , Jie Li , Zhuoqi Ma , Scott Collins , Harrison Bai , Paul Zhang , Terrance Healey , Xinbo Gao , Michael K. Atalay , Zhicheng Jiao

CLIP and BiomedCLIP are examples of vision-language foundation models and offer strong cross-modal embeddings; however, they are not optimized for fine-grained medical retrieval tasks, such as retrieving clinically relevant radiology…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Zhaohui Liang , Sivaramakrishnan Rajaraman , Niccolo Marini , Zhiyun Xue , Sameer Antani

Chest X-ray (CXR) imaging remains one of the most widely used diagnostic tools for detecting pulmonary diseases such as tuberculosis (TB) and pneumonia. Recent advances in deep learning, particularly Vision Transformers (ViTs), have shown…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Faisal Ahmed

Counterfactual image generation is a powerful tool for augmenting training data, de-biasing datasets, and modeling disease. Current approaches rely on external classifiers or regressors to increase the effectiveness of subject-level…

Recent advances in Large Language Models (LLMs) have stimulated a surge of research aimed at extending their applications to the visual domain. While these models exhibit promise in generating abstract image captions and facilitating…

Computation and Language · Computer Science 2023-10-27 Geewook Kim , Hodong Lee , Daehee Kim , Haeji Jung , Sanghee Park , Yoonsik Kim , Sangdoo Yun , Taeho Kil , Bado Lee , Seunghyun Park

Automated radiology report generation aims to generate radiology reports that contain rich, fine-grained descriptions of radiology imaging. Compared with image captioning in the natural image domain, medical images are very similar to each…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Yuhao Wang

Recent advancements in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models' advanced reasoning ability. Traditional Vision-Language Models (VLMs) perform well in visual perception tasks…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Zhiyuan Li , Dongnan Liu , Chaoyi Zhang , Heng Wang , Tengfei Xue , Weidong Cai

Contrastive Language-Image Pre-training (CLIP) has significantly improved performance in various vision-language tasks by expanding the dataset with image-text pairs obtained from websites. This paper further explores CLIP from the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Tiancheng Gu , Kaicheng Yang , Xiang An , Ziyong Feng , Dongnan Liu , Weidong Cai , Jiankang Deng

Multimodal models, such as the Contrastive Language-Image Pre-training (CLIP) model, have demonstrated remarkable success in aligning visual and linguistic representations. However, these models exhibit limitations when applied to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Hiroshi Sasaki

The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions where traditional…

Sound · Computer Science 2025-05-16 Detao Bai , Zhiheng Ma , Xihan Wei , Liefeng Bo

Foundation models leveraging vision-language pretraining have shown promise in chest X-ray (CXR) interpretation, yet their real-world performance across diverse populations and diagnostic tasks remains insufficiently evaluated. This study…