English
Related papers

Related papers: MAIRA-Seg: Enhancing Radiology Report Generation w…

200 papers

Inspired by the tremendous success of Large Language Models (LLMs), existing Radiology report generation methods attempt to leverage large models to achieve better performance. They usually adopt a Transformer to extract the visual features…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Xiao Wang , Yuehang Li , Fuling Wang , Shiao Wang , Chuanfu Li , Bo Jiang

Radiology Report Generation (R2Gen) demonstrates how Multi-modal Large Language Models (MLLMs) can automate the creation of accurate and coherent radiological reports. Existing methods often hallucinate details in text-based reports that…

Computation and Language · Computer Science 2024-07-19 Manav Nitin Kapadnis , Sohan Patnaik , Abhilash Nandy , Sourjyadip Ray , Pawan Goyal , Debdoot Sheet

Text-to-image retrieval (TIR) aims to find relevant images based on a textual query, but existing approaches are primarily based on whole-image captions and lack interpretability. Meanwhile, referring expression segmentation (RES) enables…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Li-Cheng Shen , Jih-Kang Hsieh , Wei-Hua Li , Chu-Song Chen

In this study, we propose a novel method called region-guided masked image modeling (RGMIM) for learning meaningful representations from X-ray images. Our method adopts a new masking strategy that utilizes organ mask information to identify…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Guang Li , Ren Togo , Takahiro Ogawa , Miki Haseyama

Computed tomography (CT) report generation is crucial to assist radiologists in interpreting CT volumes, which can be time-consuming and labor-intensive. Existing methods primarily only consider the global features of the entire volume,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Zhixuan Chen , Yequan Bie , Haibo Jin , Hao Chen

Automated radiology report generation holds significant potential to reduce radiologists' workload and enhance diagnostic accuracy. However, generating precise and clinically meaningful reports from chest radiographs remains challenging due…

Image and Video Processing · Electrical Eng. & Systems 2026-05-20 Md. Zihad Bin Jahangir , Muhammad Ashad Kabir , Sumaiya Akter , Israt Jahan , Minh Chau

The vision-language modeling capability of multi-modal large language models has attracted wide attention from the community. However, in medical domain, radiology report generation using vision-language models still faces significant…

Computer Vision and Pattern Recognition · Computer Science 2024-08-23 Yuhao Wang , Chao Hao , Yawen Cui , Xinqi Su , Weicheng Xie , Tao Tan , Zitong Yu

Multimodal Large Language Models (MLLMs) have shown impressive performance in vision and text tasks. However, hallucination remains a major challenge, especially in fields like healthcare where details are critical. In this work, we show…

Computation and Language · Computer Science 2025-02-24 Yun-Wei Chu , Kai Zhang , Christopher Malon , Martin Renqiang Min

Recently, Artificial Intelligence (AI)-based algorithms have revolutionized the medical image segmentation processes. Thus, the precise segmentation of organs and their lesions may contribute to an efficient diagnostics process and a more…

Neurons and Cognition · Quantitative Biology 2024-03-21 Zofia Rudnicka , Janusz Szczepanski , Agnieszka Pregowska

The applicability of current lesion segmentation models for chest X-rays (CXRs) has been limited both by a small number of target labels and the reliance on complex, expert-level text inputs, creating a barrier to practical use. To address…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Geon Choi , Hangyul Yoon , Hyunju Shin , Hyunki Park , Sang Hoon Seo , Eunho Yang , Edward Choi

Semantic segmentation of medical images is pivotal in applications like disease diagnosis and treatment planning. While deep learning has excelled in automating this task, a major hurdle is the need for numerous annotated segmentation…

Image and Video Processing · Electrical Eng. & Systems 2024-09-02 Li Zhang , Basu Jindal , Ahmed Alaa , Robert Weinreb , David Wilson , Eran Segal , James Zou , Pengtao Xie

Radiology report generation represents a significant application within medical AI, and has achieved impressive results. Concurrently, large language models (LLMs) have demonstrated remarkable performance across various domains. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Haifeng Zhao , Yufei Zhang , Leilei Ma , Shuo Xu , Dengdi Sun

Multimodal Large Language Models (MLLMs) have shown success in various general image processing tasks, yet their application in medical imaging is nascent, lacking tailored models. This study investigates the potential of MLLMs in improving…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Ling Yang , Zhanyu Wang , Zhenghao Chen , Xinyu Liang , Luping Zhou

Recent advances in reasoning-enhanced large language models (LLMs) and multimodal LLMs (MLLMs) have significantly improved performance in complex tasks, yet medical AI models often overlook the structured reasoning processes inherent in…

Artificial Intelligence · Computer Science 2025-05-22 Ziqing Fan , Cheng Liang , Chaoyi Wu , Ya Zhang , Yanfeng Wang , Weidi Xie

Chest radiography is an effective screening tool for diagnosing pulmonary diseases. In computer-aided diagnosis, extracting the relevant region of interest, i.e., isolating the lung region of each radiography image, can be an essential step…

Image and Video Processing · Electrical Eng. & Systems 2022-02-23 Hilda Azimi , Jianxing Zhang , Pengcheng Xi , Hala Asad , Ashkan Ebadi , Stephane Tremblay , Alexander Wong

While large multi-modal models (LMMs) demonstrate promising capabilities in segmentation and comprehension, they still struggle with two limitations: inaccurate segmentation and hallucinated comprehension. These challenges stem primarily…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Zhang Li , Biao Yang , Qiang Liu , Shuo Zhang , Zhiyin Ma , Liang Yin , Linger Deng , Yabo Sun , Yuliang Liu , Xiang Bai

Reliable end-to-end clinical report generation has been a longstanding goal of medical ML research. The end goal for this process is to alleviate radiologists' workloads and provide second opinions to clinicians or patients. Thus, a…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Frederic Jonske , Constantin Seibold , Osman Alperen Koras , Fin Bahnsen , Marie Bauer , Amin Dada , Hamza Kalisch , Anton Schily , Jens Kleesiek

Automatic radiology report generation can alleviate the workload for physicians and minimize regional disparities in medical resources, therefore becoming an important topic in the medical image analysis field. It is a challenging task, as…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Xinyi Wang , Grazziela Figueredo , Ruizhe Li , Wei Emma Zhang , Weitong Chen , Xin Chen

Medical image segmentation aims to identify and locate abnormal structures in medical images, such as chest radiographs, using deep neural networks. These networks require a large number of annotated images with fine-grained masks for the…

Image and Video Processing · Electrical Eng. & Systems 2024-01-17 Jiamin Chen , Xuhong Li , Yanwu Xu , Mengnan Du , Haoyi Xiong

Vision-language pretraining has advanced image-text alignment, yet progress in radiology remains constrained by the heterogeneity of clinical reports, including abbreviations, impression-only notes, and stylistic variability. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Hanbin Ko , Gihun Cho , Inhyeok Baek , Donguk Kim , Joonbeom Koo , Changi Kim , Dongheon Lee , Chang Min Park