English
Related papers

Related papers: XrayGPT: Chest Radiographs Summarization using Med…

200 papers

Generating accurate and clinically meaningful radiology reports from chest X-ray images remains a significant challenge in medical AI. While recent vision-language models achieve strong results in general radiology report generation, they…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Nikolay Nechaev , Evgeniia Przhezdzetskaia , Dmitry Umerenkov , Dmitry V. Dylov

The advent of ChatGPT and GPT-4 has captivated the world with large language models (LLMs), demonstrating exceptional performance in question-answering, summarization, and content generation. The aviation industry is characterized by an…

Computation and Language · Computer Science 2023-11-30 Liya Wang , Jason Chou , Xin Zhou , Alex Tien , Diane M Baumgartner

Since the release of ChatGPT and GPT-4, large language models (LLMs) and multimodal large language models (MLLMs) have attracted widespread attention for their exceptional capabilities in understanding, reasoning, and generation,…

Computation and Language · Computer Science 2024-12-31 Hanguang Xiao , Feizhong Zhou , Xingyue Liu , Tianqi Liu , Zhipeng Li , Xin Liu , Xiaoxuan Huang

Clinicians spend a significant amount of time reviewing medical images and transcribing their findings regarding patient diagnosis, referral and treatment in text form. Vision-language models (VLMs), which automatically interpret images and…

The success of Multimodal Large Language Models (MLLMs) in the medical auxiliary field shows great potential, allowing patients to engage in conversations using physiological signal data. However, general MLLMs perform poorly in cardiac…

Signal Processing · Electrical Eng. & Systems 2025-04-17 Yubao Zhao , Jiaju Kang , Tian Zhang , Puyu Han , Tong Chen

Following the impressive development of LLMs, vision-language alignment in LLMs is actively being researched to enable multimodal reasoning and visual IO. This direction of research is particularly relevant to medical imaging because…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Suhyeon Lee , Won Jun Kim , Jinho Chang , Jong Chul Ye

The scarcity of well-annotated diverse medical images is a major hurdle for developing reliable AI models in healthcare. Substantial technical advances have been made in generative foundation models for natural images. Here we develop…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Yuanfeng Ji , Dan Lin , Xiyue Wang , Lu Zhang , Wenhui Zhou , Chongjian Ge , Ruihang Chu , Xiaoli Yang , Junhan Zhao , Junsong Chen , Xiangde Luo , Sen Yang , Jin Fang , Ping Luo , Ruijiang Li

Chest X-ray imaging is crucial for diagnosing pulmonary and cardiac diseases, yet its interpretation demands extensive clinical experience and suffers from inter-observer variability. While deep learning models offer high diagnostic…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Chee Ng , Liliang Sun , Shaoqing Tang

Recent vision-language models (VLMs) have shown strong generalization and multimodal reasoning abilities in natural domains. However, their application to medical diagnosis remains limited by the lack of comprehensive and structured…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Sheng Lu , Hao Chen , Rui Yin , Juyan Ba , Yu Zhang , Yuanzhe Li

Large vision language models (VLMs) have progressed incredibly from research to applicability for general-purpose use cases. LLaVA-Med, a pioneering large language and vision assistant for biomedicine, can perform multi-modal biomedical…

State-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLaVA-Med and BioMedGPT, primarily depend on scaling model size and data volume, with training driven largely by autoregressive objectives. However, we reveal that this approach…

With the availability of large-scale, comprehensive, and general-purpose vision-language (VL) datasets such as MSCOCO, vision-language pre-training (VLP) has become an active area of research and proven to be effective for various VL tasks…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Li Xu , Bo Liu , Ameer Hamza Khan , Lu Fan , Xiao-Ming Wu

The high incidence and mortality rates associated with respiratory diseases underscores the importance of early screening. Machine learning models can automate clinical consultations and auscultation, offering vital support in this area.…

Machine Learning · Computer Science 2024-10-10 Yuwei Zhang , Tong Xia , Aaqib Saeed , Cecilia Mascolo

Oversight AI is an emerging concept in radiology where the AI forms a symbiosis with radiologists by continuously supporting radiologists in their decision-making. Recent advances in vision-language models sheds a light on the long-standing…

Image and Video Processing · Electrical Eng. & Systems 2023-04-13 Sangjoon Park , Eun Sun Lee , Kyung Sook Shin , Jeong Eun Lee , Jong Chul Ye

Chest X-ray interpretation is one of the most frequently performed diagnostic tasks in medicine and a primary target for AI development, yet current vision-language models are primarily trained on datasets of paired images and reports, not…

Image-to-text radiology report generation aims to automatically produce radiology reports that describe the findings in medical images. Most existing methods focus solely on the image data, disregarding the other patient information…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Nurbanu Aksoy , Serge Sharoff , Selcuk Baser , Nishant Ravikumar , Alejandro F Frangi

Our paper focuses on automating the generation of medical reports from chest X-ray image inputs, a critical yet time-consuming task for radiologists. Unlike existing medical re-port generation efforts that tend to produce human-readable…

Computation and Language · Computer Science 2021-08-30 Hoang T. N. Nguyen , Dong Nie , Taivanbat Badamdorj , Yujie Liu , Yingying Zhu , Jason Truong , Li Cheng

Medical image-language pre-training aims to align medical images with clinically relevant text to improve model performance on various downstream tasks. However, existing models often struggle with the variability and ambiguity inherent in…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Shreyank N Gowda , Ruichi Zhang , Xiao Gu , Ying Weng , Lu Yang

Analyzing vast textual data and summarizing key information from electronic health records imposes a substantial burden on how clinicians allocate their time. Although large language models (LLMs) have shown promise in natural language…

This paper introduces MiniGPT4-Video, a multimodal Large Language Model (LLM) designed specifically for video understanding. The model is capable of processing both temporal visual and textual data, making it adept at understanding the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-05 Kirolos Ataallah , Xiaoqian Shen , Eslam Abdelrahman , Essam Sleiman , Deyao Zhu , Jian Ding , Mohamed Elhoseiny