中文
相关论文

相关论文: GPT-4V Cannot Generate Radiology Reports Yet

200 篇论文

The automation of chest X-ray reporting has garnered significant interest due to the time-consuming nature of the task. However, the clinical accuracy of free-text reports has proven challenging to quantify using natural language processing…

计算机视觉与模式识别 · 计算机科学 2023-05-03 Matthias Keicher , Kamilia Zaripova , Tobias Czempiel , Kristina Mach , Ashkan Khakzar , Nassir Navab

Recent advances in deep learning have enabled researchers to explore tasks at the intersection of computer vision and natural language processing, such as image captioning, visual question answering, visual dialogue, and visual language…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Sonit Singh

Radiology reports are crucial for planning treatment strategies and facilitating effective doctor-patient communication. However, the manual creation of these reports places a significant burden on radiologists. While automatic radiology…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Qiguang Miao , Kang Liu , Zhuoqi Ma , Yunan Li , Xiaolu Kang , Ruixuan Liu , Tianyi Liu , Kun Xie , Zhicheng Jiao

Many real-world tasks require an agent to reason jointly over text and visual objects, (e.g., navigating in public spaces), which we refer to as context-sensitive text-rich visual reasoning. Specifically, these tasks require an…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Rohan Wadhawan , Hritik Bansal , Kai-Wei Chang , Nanyun Peng

Automatic medical report generation from chest X-ray images is one possibility for assisting doctors to reduce their workload. However, the different patterns and data distribution of normal and abnormal cases can bias machine learning…

计算机视觉与模式识别 · 计算机科学 2019-08-07 Philipp Harzig , Yan-Ying Chen , Francine Chen , Rainer Lienhart

AI-assisted report generation offers the opportunity to reduce radiologists' workload stemming from expanded screening guidelines, complex cases and workforce shortages, while maintaining diagnostic accuracy. In addition to describing…

The release of ChatGPT, a language model capable of generating text that appears human-like and authentic, has gained significant attention beyond the research community. We expect that the convincing performance of ChatGPT incentivizes…

Vision-language models (VLMs) have shown strong promise for medical image analysis, but most remain opaque, offering predictions without the transparent, stepwise reasoning clinicians rely on. We present a framework that brings…

Radiology Report Generation (RRG) has advanced considerably with the development of multimodal generative models. Despite the progress, the field still faces significant challenges in evaluation, as existing metrics lack robustness and…

计算与语言 · 计算机科学 2025-05-20 Kun Zhao , Chenghao Xiao , Sixing Yan , Haoteng Tang , William K. Cheung , Noura Al Moubayed , Liang Zhan , Chenghua Lin

Harnessing logical reasoning ability is a comprehensive natural language understanding endeavor. With the release of Generative Pretrained Transformer 4 (GPT-4), highlighted as "advanced" at reasoning tasks, we are eager to learn the GPT-4…

计算与语言 · 计算机科学 2023-05-08 Hanmeng Liu , Ruoxi Ning , Zhiyang Teng , Jian Liu , Qiji Zhou , Yue Zhang

With the emergence of large-scale vision-language models, realistic radiology reports may be generated using only medical images as input guided by simple prompts. However, their practical utility has been limited due to the factual errors…

计算机视觉与模式识别 · 计算机科学 2024-12-04 R. Mahmood , K. C. L. Wong , D. M. Reyes , N. D'Souza , L. Shi , J. Wu , P. Kaviani , M. Kalra , G. Wang , P. Yan , T. Syeda-Mahmood

We present a novel approach to Chest X-ray (CXR) Visual Question Answering (VQA), addressing both single-image image-difference questions. Single-image questions focus on abnormalities within a specific CXR ("What abnormalities are seen in…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Francesco Dalla Serra , Patrick Schrempf , Chaoyang Wang , Zaiqiao Meng , Fani Deligianni , Alison Q. O'Neil

Recent vision-language models (VLMs) have shown strong generalization and multimodal reasoning abilities in natural domains. However, their application to medical diagnosis remains limited by the lack of comprehensive and structured…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Sheng Lu , Hao Chen , Rui Yin , Juyan Ba , Yu Zhang , Yuanzhe Li

Large vision-language models (LVMs) hold a great promise for automating medical report generation, potentially reducing the burden of manual reporting. State-of-the-art (SOTA) research fine-tunes general LVMs with medical data to align…

计算与语言 · 计算机科学 2025-04-07 Hao Wang , Shuchang Ye , Jinghao Lin , Usman Naseem , Jinman Kim

Automatic radiology report generation has been an attracting research problem towards computer-aided diagnosis to alleviate the workload of doctors in recent years. Deep learning techniques for natural image captioning are successfully…

计算机视觉与模式识别 · 计算机科学 2020-02-20 Yixiao Zhang , Xiaosong Wang , Ziyue Xu , Qihang Yu , Alan Yuille , Daguang Xu

Multimodal foundation models have shown compelling but conflicting performance in medical image interpretation. However, the mechanisms by which these models integrate and prioritize different data modalities, including images and text,…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Thomas Buckley , James A. Diao , Pranav Rajpurkar , Adam Rodman , Arjun K. Manrai

The remote sensing image intelligence understanding model is undergoing a new profound paradigm shift which has been promoted by multi-modal large language model (MLLM), i.e. from the paradigm learning a domain model (LaDM) shifts to…

计算机视觉与模式识别 · 计算机科学 2024-06-19 Linrui Xu , Ling Zhao , Wang Guo , Qiujun Li , Kewang Long , Kaiqi Zou , Yuhan Wang , Haifeng Li

Scoring student-drawn models is time-consuming. Recently released GPT-4V provides a unique opportunity to advance scientific modeling practices by leveraging the powerful image processing capability. To test this ability specifically for…

人工智能 · 计算机科学 2023-12-29 Gyeong-Geon Lee , Xiaoming Zhai

As artificial intelligence (AI) becomes increasingly central to healthcare, the demand for explainable and trustworthy models is paramount. Current report generation systems for chest X-rays (CXR) often lack mechanisms for validating…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Sayeh Gholipour Picha , Dawood Al Chanti , Alice Caplier

Recent developments in multimodal methodologies have marked the beginning of an exciting era for models adept at processing diverse data types, encompassing text, audio, and visual content. Models like GPT-4V, which merge computer vision…