English
Related papers

Related papers: S2D-ALIGN: Shallow-to-Deep Auxiliary Learning for …

200 papers

Radiology report generation (RRG) models typically focus on individual exams, often overlooking the integration of historical visual or textual data, which is crucial for patient follow-ups. Traditional methods usually struggle with long…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Tengfei Liu , Jiapu Wang , Yongli Hu , Mingjie Li , Junfei Yi , Xiaojun Chang , Junbin Gao , Baocai Yin

Radiology report generation requires advanced medical image analysis, effective temporal reasoning, and accurate text generation. Although recent innovations, particularly multimodal large language models, have shown improved performance,…

Computation and Language · Computer Science 2025-11-11 Kai Zhang , Christopher Malon , Lichao Sun , Martin Renqiang Min

Spatial transcriptomics (ST) provides high-resolution pathological images and whole-transcriptomic expression profiles at individual spots across whole-slide scales. This setting makes it an ideal data source to develop multimodal…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Yuxiang Lin , Ling Luo , Ying Chen , Xushi Zhang , Zihui Wang , Wenxian Yang , Mengsha Tong , Rongshan Yu

Radiology Report Generation (RRG) is an important research topic for relieving radiologist' heavy workload. Existing RRG models mainly rely on supervised fine-tuning (SFT) based on different model architectures using data pairs of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Ting Xiao , Lei Shi , Yang Zhang , HaoFeng Yang , Zhe Wang , Chenjia Bai

Medical image segmentation driven by free-text clinical instructions is a critical frontier in computer-aided diagnosis. However, existing multimodal and foundation models struggle with the semantic ambiguity of clinical reports and fail to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Chenyu Xue , Yiran Liu , Mian Zhou , Jionglong Su , Zhixiang Lu

We propose Retrieval Augmented Generation (RAG) as an approach for automated radiology report writing that leverages multimodally aligned embeddings from a contrastively pretrained vision language model for retrieval of relevant candidate…

Computation and Language · Computer Science 2023-05-08 Mercy Ranjit , Gopinath Ganapathy , Ranjit Manuel , Tanuja Ganu

The automatic generation of radiology reports has the potential to assist radiologists in the time-consuming task of report writing. Existing methods generate the full report from image-level features, failing to explicitly focus on…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Tim Tanida , Philip Müller , Georgios Kaissis , Daniel Rueckert

Developing 3D vision-language models with robust clinical reasoning remains a challenge due to the inherent complexity of volumetric medical imaging, the tendency of models to overfit superficial report patterns, and the lack of…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Haoran Lai , Zihang Jiang , Kun Zhang , Qingsong Yao , Rongsheng Wang , Zhiyang He , Xiaodong Tao , Wei Wei , Shaohua Kevin Zhou

The application of Vision-Language Models (VLMs) in medicine is critically hampered by the scarcity of high-quality, expert-annotated data. Supervised Fine-Tuning (SFT) on existing datasets often leads to poor generalization on unseen…

Machine Learning · Computer Science 2025-12-09 Weihai Zhi , Jiayan Guo , Shangyang Li

Automatically generated reports from medical images promise to improve the workflow of radiologists. Existing methods consider an image-to-report modeling task by directly generating a fully-fledged report from an image. However, this…

Automated radiology report generation aims at automatically generating a detailed description of medical images, which can greatly alleviate the workload of radiologists and provide better medical services to remote areas. Most existing…

Computer Vision and Pattern Recognition · Computer Science 2022-11-22 Yuhao Wang , Kai Wang , Xiaohong Liu , Tianrun Gao , Jingyue Zhang , Guangyu Wang

Automatic Medical Imaging Narrative generation aims to alleviate the workload of radiologists by producing accurate clinical descriptions directly from radiological images. However, the subtle visual nuances and domain-specific terminology…

Computer Vision and Pattern Recognition · Computer Science 2024-09-09 Kai Shu , Yuzhuo Jia , Ziyang Zhang , Jiechao Gao

Radiology reporting is a complex task requiring detailed medical image understanding and precise language generation, for which generative multimodal models offer a promising solution. However, to impact clinical practice, models must…

Large language models (LLMs) often generate outdated or inaccurate information based on static training datasets. Retrieval-augmented generation (RAG) mitigates this by integrating outside data sources. While previous RAG systems used…

Medical report generation demands automatic creation of coherent and precise descriptions for medical images. However, the scarcity of labelled medical image-report pairs poses formidable challenges in developing large-scale neural networks…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Shibin Wu , Bang Yang , Zhiyu Ye , Haoqian Wang , Hairong Zheng , Tong Zhang

Advancements in generative Artificial Intelligence (AI) hold great promise for automating radiology workflows, yet challenges in interpretability and reliability hinder clinical adoption. This paper presents an automated radiology report…

Medical image interpretation is central to most clinical applications such as disease diagnosis, treatment planning, and prognostication. In clinical practice, radiologists examine medical images and manually compile their findings into…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Nurbanu Aksoy , Nishant Ravikumar , Alejandro F Frangi

In the current paradigm of image captioning, deep learning models are trained to generate text from image embeddings of latent features. We challenge the assumption that fine-tuning of large, bespoke models is required to improve model…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Steven Song , Anirudh Subramanyam , Irene Madejski , Robert L. Grossman

Beyond the common difficulties faced in the natural image captioning, medical report generation specifically requires the model to describe a medical image with a fine-grained and semantic-coherence paragraph that should satisfy both…

Computer Vision and Pattern Recognition · Computer Science 2020-06-09 Mingjie Li , Fuyu Wang , Xiaojun Chang , Xiaodan Liang

Recent advances in automated radiology report generation from chest X-rays using deep learning algorithms have the potential to significantly reduce the arduous workload of radiologists. However, due to the inherent massive data bias in…

Computer Vision and Pattern Recognition · Computer Science 2025-07-16 Zeyi Hou , Zeqiang Wei , Ruixin Yan , Ning Lang , Xiuzhuang Zhou