English
Related papers

Related papers: RoentGen: Vision-Language Foundation Model for Che…

200 papers

Phrase grounding, i.e., mapping natural language phrases to specific image regions, holds significant potential for disease localization in medical imaging through clinical reports. While current state-of-the-art methods rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Felix Nützel , Mischa Dombrowski , Bernhard Kainz

Chest X-ray (CXR) imaging is one of the most widely used diagnostic modalities in clinical practice, encompassing a broad spectrum of diagnostic tasks. Recent advancements have seen the extensive application of reasoning-based multimodal…

Automatic radiology report generation is a promising application of multimodal deep learning, aiming to reduce reporting workload and improve consistency. However, current state-of-the-art (SOTA) systems - such as Multimodal AI for…

We present Imagen, a text-to-image diffusion model with an unprecedented degree of photorealism and a deep level of language understanding. Imagen builds on the power of large transformer language models in understanding text and hinges on…

In clinics, a radiology report is crucial for guiding a patient's treatment. However, writing radiology reports is a heavy burden for radiologists. To this end, we present an automatic, multi-modal approach for report generation from a…

Image and Video Processing · Electrical Eng. & Systems 2022-06-02 Shuxin Yang , Xian Wu , Shen Ge , S. Kevin Zhou , Li Xiao

Text to image latent diffusion models have recently advanced medical image synthesis, but applications to 3D CT generation remain limited. Existing approaches rely on simplified prompts, neglecting the rich semantic detail in full radiology…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Sina Amirrajab , Zohaib Salahuddin , Sheng Kuang , Henry C. Woodruff , Philippe Lambin

Chest X-rays (CXRs) are among the most frequently performed imaging examinations worldwide, yet rising imaging volumes increase radiologist workload and the risk of diagnostic errors. Although artificial intelligence (AI) systems have shown…

Multimodal large models have shown great potential in automating pathology image analysis. However, current multimodal models for gastrointestinal pathology are constrained by both data quality and reasoning transparency: pervasive noise…

Image and Video Processing · Electrical Eng. & Systems 2025-07-25 Minxi Ouyang , Lianghui Zhu , Yaqing Bao , Qiang Huang , Jingli Ouyang , Tian Guan , Xitong Ling , Jiawen Li , Song Duan , Wenbin Dai , Li Zheng , Xuemei Zhang , Yonghong He

Radiology narrative reports often describe characteristics of a patient's disease, including its location, size, and shape. Motivated by the recent success of multimodal learning, we hypothesized that this descriptive text could guide…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Zachary Huemann , Xin Tie , Junjie Hu , Tyler J. Bradshaw

Automating radiology report generation can ease the reporting workload for radiologists. However, existing works focus mainly on the chest area due to the limited availability of public datasets for other regions. Besides, they often rely…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Qi Chen , Yutong Xie , Biao Wu , Xiaomin Chen , James Ang , Minh-Son To , Xiaojun Chang , Qi Wu

Despite the progress in automatic detection of radiologic findings from chest X-ray (CXR) images in recent years, a quantitative evaluation of the explainability of these models is hampered by the lack of locally labeled datasets for…

Generating radiology reports is time-consuming and requires extensive expertise in practice. Therefore, reliable automatic radiology report generation is highly desired to alleviate the workload. Although deep learning techniques have been…

Image and Video Processing · Electrical Eng. & Systems 2019-07-24 Jianbo Yuan , Haofu Liao , Rui Luo , Jiebo Luo

While diffusion models have significantly advanced the quality of image generation their capability to accurately and coherently render text within these images remains a substantial challenge. Conventional diffusion-based methods for scene…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Qilong Zhangli , Jindong Jiang , Di Liu , Licheng Yu , Xiaoliang Dai , Ankit Ramchandani , Guan Pang , Dimitris N. Metaxas , Praveen Krishnan

Generative models are becoming popular for the synthesis of medical images. Recently, neural diffusion models have demonstrated the potential to generate photo-realistic images of objects. However, their potential to generate medical images…

Image and Video Processing · Electrical Eng. & Systems 2022-11-03 Hazrat Ali , Shafaq Murad , Zubair Shah

Vision-language models have been extensively explored across a wide range of tasks, achieving satisfactory performance; however, their application in medical imaging remains underexplored. In this work, we propose a unified framework -…

Image and Video Processing · Electrical Eng. & Systems 2024-07-18 Khai Le-Duc , Ryan Zhang , Ngoc Son Nguyen , Tan-Hanh Pham , Anh Dao , Ba Hung Ngo , Anh Totti Nguyen , Truong-Son Hy

Textual image generation spans diverse fields like advertising, education, product packaging, social media, information visualization, and branding. Despite recent strides in language-guided image synthesis using diffusion models, current…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Shubham Paliwal , Arushi Jain , Monika Sharma , Vikram Jamwal , Lovekesh Vig

The availability of large-scale chest X-ray datasets is a requirement for developing well-performing deep learning-based algorithms in thoracic abnormality detection and classification. However, biometric identifiers in chest radiographs…

Image and Video Processing · Electrical Eng. & Systems 2024-10-28 Kai Packhäuser , Lukas Folle , Florian Thamm , Andreas Maier

Vision-language models (VLMs) have shown strong promise for medical image analysis, but most remain opaque, offering predictions without the transparent, stepwise reasoning clinicians rely on. We present a framework that brings…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Andriy Myronenko , Dong Yang , Baris Turkbey , Mariam Aboian , Sena Azamat , Esra Akcicek , Hongxu Yin , Pavlo Molchanov , Marc Edgar , Yufan He , Pengfei Guo , Yucheng Tang , Daguang Xu

Large vision-language models (VLMs) have evolved from general-purpose applications to specialized use cases such as in the clinical domain, demonstrating potential for decision support in radiology. One promising application is assisting…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Zhifan Jiang , Dong Yang , Vishwesh Nath , Abhijeet Parida , Nishad P. Kulkarni , Ziyue Xu , Daguang Xu , Syed Muhammad Anwar , Holger R. Roth , Marius George Linguraru

To effectively train medical students to become qualified radiologists, a large number of X-ray images collected from patients with diverse medical conditions are needed. However, due to data privacy concerns, such images are typically…

Image and Video Processing · Electrical Eng. & Systems 2020-06-19 Xingyi Yang , Nandiraju Gireesh , Eric Xing , Pengtao Xie