English
Related papers

Related papers: GPT-4V Cannot Generate Radiology Reports Yet

200 papers

The automation of chest X-ray reporting has garnered significant interest due to the time-consuming nature of the task. However, the clinical accuracy of free-text reports has proven challenging to quantify using natural language processing…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Matthias Keicher , Kamilia Zaripova , Tobias Czempiel , Kristina Mach , Ashkan Khakzar , Nassir Navab

Recent advances in deep learning have enabled researchers to explore tasks at the intersection of computer vision and natural language processing, such as image captioning, visual question answering, visual dialogue, and visual language…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Sonit Singh

Radiology reports are crucial for planning treatment strategies and facilitating effective doctor-patient communication. However, the manual creation of these reports places a significant burden on radiologists. While automatic radiology…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Qiguang Miao , Kang Liu , Zhuoqi Ma , Yunan Li , Xiaolu Kang , Ruixuan Liu , Tianyi Liu , Kun Xie , Zhicheng Jiao

Many real-world tasks require an agent to reason jointly over text and visual objects, (e.g., navigating in public spaces), which we refer to as context-sensitive text-rich visual reasoning. Specifically, these tasks require an…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Rohan Wadhawan , Hritik Bansal , Kai-Wei Chang , Nanyun Peng

Automatic medical report generation from chest X-ray images is one possibility for assisting doctors to reduce their workload. However, the different patterns and data distribution of normal and abnormal cases can bias machine learning…

Computer Vision and Pattern Recognition · Computer Science 2019-08-07 Philipp Harzig , Yan-Ying Chen , Francine Chen , Rainer Lienhart

AI-assisted report generation offers the opportunity to reduce radiologists' workload stemming from expanded screening guidelines, complex cases and workforce shortages, while maintaining diagnostic accuracy. In addition to describing…

The release of ChatGPT, a language model capable of generating text that appears human-like and authentic, has gained significant attention beyond the research community. We expect that the convincing performance of ChatGPT incentivizes…

Vision-language models (VLMs) have shown strong promise for medical image analysis, but most remain opaque, offering predictions without the transparent, stepwise reasoning clinicians rely on. We present a framework that brings…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Andriy Myronenko , Dong Yang , Baris Turkbey , Mariam Aboian , Sena Azamat , Esra Akcicek , Hongxu Yin , Pavlo Molchanov , Marc Edgar , Yufan He , Pengfei Guo , Yucheng Tang , Daguang Xu

Radiology Report Generation (RRG) has advanced considerably with the development of multimodal generative models. Despite the progress, the field still faces significant challenges in evaluation, as existing metrics lack robustness and…

Computation and Language · Computer Science 2025-05-20 Kun Zhao , Chenghao Xiao , Sixing Yan , Haoteng Tang , William K. Cheung , Noura Al Moubayed , Liang Zhan , Chenghua Lin

Harnessing logical reasoning ability is a comprehensive natural language understanding endeavor. With the release of Generative Pretrained Transformer 4 (GPT-4), highlighted as "advanced" at reasoning tasks, we are eager to learn the GPT-4…

Computation and Language · Computer Science 2023-05-08 Hanmeng Liu , Ruoxi Ning , Zhiyang Teng , Jian Liu , Qiji Zhou , Yue Zhang

With the emergence of large-scale vision-language models, realistic radiology reports may be generated using only medical images as input guided by simple prompts. However, their practical utility has been limited due to the factual errors…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 R. Mahmood , K. C. L. Wong , D. M. Reyes , N. D'Souza , L. Shi , J. Wu , P. Kaviani , M. Kalra , G. Wang , P. Yan , T. Syeda-Mahmood

We present a novel approach to Chest X-ray (CXR) Visual Question Answering (VQA), addressing both single-image image-difference questions. Single-image questions focus on abnormalities within a specific CXR ("What abnormalities are seen in…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Francesco Dalla Serra , Patrick Schrempf , Chaoyang Wang , Zaiqiao Meng , Fani Deligianni , Alison Q. O'Neil

Recent vision-language models (VLMs) have shown strong generalization and multimodal reasoning abilities in natural domains. However, their application to medical diagnosis remains limited by the lack of comprehensive and structured…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Sheng Lu , Hao Chen , Rui Yin , Juyan Ba , Yu Zhang , Yuanzhe Li

Large vision-language models (LVMs) hold a great promise for automating medical report generation, potentially reducing the burden of manual reporting. State-of-the-art (SOTA) research fine-tunes general LVMs with medical data to align…

Computation and Language · Computer Science 2025-04-07 Hao Wang , Shuchang Ye , Jinghao Lin , Usman Naseem , Jinman Kim

Automatic radiology report generation has been an attracting research problem towards computer-aided diagnosis to alleviate the workload of doctors in recent years. Deep learning techniques for natural image captioning are successfully…

Computer Vision and Pattern Recognition · Computer Science 2020-02-20 Yixiao Zhang , Xiaosong Wang , Ziyue Xu , Qihang Yu , Alan Yuille , Daguang Xu

Multimodal foundation models have shown compelling but conflicting performance in medical image interpretation. However, the mechanisms by which these models integrate and prioritize different data modalities, including images and text,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Thomas Buckley , James A. Diao , Pranav Rajpurkar , Adam Rodman , Arjun K. Manrai

The remote sensing image intelligence understanding model is undergoing a new profound paradigm shift which has been promoted by multi-modal large language model (MLLM), i.e. from the paradigm learning a domain model (LaDM) shifts to…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Linrui Xu , Ling Zhao , Wang Guo , Qiujun Li , Kewang Long , Kaiqi Zou , Yuhan Wang , Haifeng Li

Scoring student-drawn models is time-consuming. Recently released GPT-4V provides a unique opportunity to advance scientific modeling practices by leveraging the powerful image processing capability. To test this ability specifically for…

Artificial Intelligence · Computer Science 2023-12-29 Gyeong-Geon Lee , Xiaoming Zhai

As artificial intelligence (AI) becomes increasingly central to healthcare, the demand for explainable and trustworthy models is paramount. Current report generation systems for chest X-rays (CXR) often lack mechanisms for validating…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Sayeh Gholipour Picha , Dawood Al Chanti , Alice Caplier

Recent developments in multimodal methodologies have marked the beginning of an exciting era for models adept at processing diverse data types, encompassing text, audio, and visual content. Models like GPT-4V, which merge computer vision…

Computation and Language · Computer Science 2024-11-15 Xiang Zhang , Senyu Li , Ning Shi , Bradley Hauer , Zijun Wu , Grzegorz Kondrak , Muhammad Abdul-Mageed , Laks V. S. Lakshmanan