English
Related papers

Related papers: Rad-ReStruct: A Novel VQA Benchmark and Method for…

200 papers

Visual question answering requires a system to provide an accurate natural language answer given an image and a natural language question. However, it is widely recognized that previous generic VQA methods often exhibit a tendency to…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Jie Ma , Pinghui Wang , Dechen Kong , Zewei Wang , Jun Liu , Hongbin Pei , Junzhou Zhao

Most natural language tasks in the radiology domain use language models pre-trained on biomedical corpus. There are few pretrained language models trained specifically for radiology, and fewer still that have been trained in a low data…

Computation and Language · Computer Science 2023-06-06 Rikhiya Ghosh , Sanjeev Kumar Karn , Manuela Daniela Danu , Larisa Micu , Ramya Vunikili , Oladimeji Farri

In recent years, deep learning techniques have shown significant potential for improving video quality assessment (VQA), achieving higher correlation with subjective opinions compared to conventional approaches. However, the development of…

Image and Video Processing · Electrical Eng. & Systems 2024-06-11 Chen Feng , Duolikun Danier , Fan Zhang , David Bull

With the emergence of large-scale vision-language models, realistic radiology reports may be generated using only medical images as input guided by simple prompts. However, their practical utility has been limited due to the factual errors…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 R. Mahmood , K. C. L. Wong , D. M. Reyes , N. D'Souza , L. Shi , J. Wu , P. Kaviani , M. Kalra , G. Wang , P. Yan , T. Syeda-Mahmood

Neural image-to-text radiology report generation systems offer the potential to improve radiology reporting by reducing the repetitive process of report drafting and identifying possible medical errors. These systems have achieved promising…

Computation and Language · Computer Science 2022-10-25 Jean-Benoit Delbrouck , Pierre Chambon , Christian Bluethgen , Emily Tsai , Omar Almusa , Curtis P. Langlotz

The intersection of medical Visual Question Answering (Med-VQA) is a challenging research topic with advantages including patient engagement and clinical expert involvement for second opinions. However, existing Med-VQA methods based on…

Machine Learning · Computer Science 2024-06-24 Lin Fan , Xun Gong , Cenyang Zheng , Yafei Ou

Radiology report generation (RRG) aims to describe automatically a radiology image with human-like language and could potentially support the work of radiologists, reducing the burden of manual reporting. Previous approaches often adopt an…

Computer Vision and Pattern Recognition · Computer Science 2022-07-12 Jun Wang , Abhir Bhalerao , Yulan He

The intricate and multifaceted nature of vision language model (VLM) development, adaptation, and application necessitates the establishment of clear and standardized reporting protocols, particularly within the high-stakes context of…

Computers and Society · Computer Science 2025-05-15 Amara Tariq , Rimita Lahiri , Charles Kahn , Imon Banerjee

Visual Question Answering (VQA) focuses on providing answers to natural language questions by utilizing information from images. Although cutting-edge multimodal large language models (MLLMs) such as GPT-4o achieve strong performance on VQA…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Zhengxuan Zhang , Yin Wu , Yuyu Luo , Nan Tang

Extracting meaningful insights from large and complex datasets poses significant challenges, particularly in ensuring the accuracy and relevance of retrieved information. Traditional data retrieval methods such as sequential search and…

Information Retrieval · Computer Science 2024-09-27 Zahra Sepasdar , Sushant Gautam , Cise Midoglu , Michael A. Riegler , Pål Halvorsen

Automatic report generation has arisen as a significant research area in computer-aided diagnosis, aiming to alleviate the burden on clinicians by generating reports automatically based on medical images. In this work, we propose a novel…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Jun Li , Tongkun Su , Baoliang Zhao , Faqin Lv , Qiong Wang , Nassir Navab , Ying Hu , Zhongliang Jiang

Radiology report generation aims to automatically provide clinically meaningful descriptions of radiology images such as MRI and X-ray. Although great success has been achieved in natural scene image captioning tasks, radiology report…

Computer Vision and Pattern Recognition · Computer Science 2023-09-01 Jun Wang , Lixing Zhu , Abhir Bhalerao , Yulan He

Various industries have produced a large number of documents such as industrial plans, technical guidelines, and regulations that are structurally complex and content-wise fragmented. This poses significant challenges for experts and…

Artificial Intelligence · Computer Science 2025-05-27 Hongjia Wu , Hongxin Zhang , Wei Chen , Jiazhi Xia

Radiology report generation (RRG) has attracted significant attention due to its potential to reduce the workload of radiologists. Current RRG approaches are still unsatisfactory against clinical standards. This paper introduces a novel RRG…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Zijian Zhou , Miaojing Shi , Meng Wei , Oluwatosin Alabi , Zijie Yue , Tom Vercauteren

The way we analyse clinical texts has undergone major changes over the last years. The introduction of language models such as BERT led to adaptations for the (bio)medical domain like PubMedBERT and ClinicalBERT. These models rely on large…

Computation and Language · Computer Science 2023-09-15 Tom van Sonsbeek , Xiantong Zhen , Marcel Worring

We present ReXVQA, the largest and most comprehensive benchmark for visual question answering (VQA) in chest radiology, comprising approximately 696,000 questions paired with 160,000 chest X-rays studies across training, validation, and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Ankit Pal , Jung-Oh Lee , Xiaoman Zhang , Malaikannan Sankarasubbu , Seunghyeon Roh , Won Jung Kim , Meesun Lee , Pranav Rajpurkar

Large language models (LLMs), including zero-shot and few-shot paradigms, have shown promising capabilities in clinical text generation. However, real-world applications face two key challenges: (1) patient data is highly unstructured,…

Computation and Language · Computer Science 2025-07-10 Garapati Keerthana , Manik Gupta

A radiology report comprises presentation-style vocabulary, which ensures clarity and organization, and factual vocabulary, which provides accurate and objective descriptions based on observable findings. While manually writing these…

Image and Video Processing · Electrical Eng. & Systems 2026-05-28 Kang Liu , Zhuoqi Ma , Mengmeng Liu , Zhicheng Jiao , Xiaolu Kang , Qiguang Miao , Kun Xie

Large multimodal models (LMMs) have achieved impressive progress in vision-language understanding, yet they face limitations in real-world applications requiring complex reasoning over a large number of images. Existing benchmarks for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Jun Chen , Dannong Xu , Junjie Fei , Chun-Mei Feng , Mohamed Elhoseiny

Purpose: Federated training is often hindered by heterogeneous datasets due to divergent data storage options, inconsistent naming schemes, varied annotation procedures, and disparities in label quality. This is particularly evident in the…