English
Related papers

Related papers: Generalised Medical Phrase Grounding

200 papers

Automated radiology report generation is essential in clinical practice. However, diagnosing radiological images typically requires physicians 5-10 minutes, resulting in a waste of valuable healthcare resources. Existing studies have not…

Multimedia · Computer Science 2025-09-16 Jing Xiao , Hongfei Liu , Ruiqi Dong , Jimin Liu , Haoyong Yu

Automatic radiology report generation is essential to computer-aided diagnosis. Through the success of image captioning, medical report generation has been achievable. However, the lack of annotated disease labels is still the bottleneck of…

Computation and Language · Computer Science 2022-06-22 Jun Li , Shibo Li , Ying Hu , Huiren Tao

Visual grounding is essential for precise perception and reasoning in multimodal large language models (MLLMs), especially in medical imaging domains. While existing medical visual grounding benchmarks primarily focus on single-image…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Jingkun Yue , Siqi Zhang , Zinan Jia , Huihuan Xu , Zongbo Han , Xiaohong Liu , Guangyu Wang

Radiologists are tasked with interpreting a large number of images in a daily base, with the responsibility of generating corresponding reports. This demanding workload elevates the risk of human error, potentially leading to treatment…

Image and Video Processing · Electrical Eng. & Systems 2024-07-31 Jiayu Lei , Xiaoman Zhang , Chaoyi Wu , Lisong Dai , Ya Zhang , Yanyong Zhang , Yanfeng Wang , Weidi Xie , Yuehua Li

In recent years, accurately and quickly deploying medical large language models (LLMs) has become a trend. Among these, retrieval-augmented generation (RAG) has garnered attention due to rapid deployment and privacy protection. However, the…

Computation and Language · Computer Science 2025-08-06 Penglei Sun , Yixiang Chen , Xiang Li , Xiaowen Chu

Visual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) based on the given expression text. This task simulates the common referential relationships…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Linhui Xiao , Xiaoshan Yang , Xiangyuan Lan , Yaowei Wang , Changsheng Xu

Medical Image Grounding (MIG), which involves localizing specific regions in medical images based on textual descriptions, requires models to not only perceive regions but also deduce spatial relationships of these regions. Existing…

Machine Learning · Computer Science 2025-07-08 Huihui Xu , Yuanpeng Nie , Hualiang Wang , Ying Chen , Wei Li , Junzhi Ning , Lihao Liu , Hongqiu Wang , Lei Zhu , Jiyao Liu , Xiaomeng Li , Junjun He

Question answering models struggle to generalize to novel compositions of training patterns, such to longer sequences or more complex test structures. Current end-to-end models learn a flat input embedding which can lose input syntax…

Computation and Language · Computer Science 2021-11-08 Yu Gai , Paras Jain , Wendi Zhang , Joseph E. Gonzalez , Dawn Song , Ion Stoica

Robots collaborating with humans must convert natural language goals into actionable, physically grounded decisions. For example, executing a command such as "go two meters to the right of the fridge" requires grounding semantic references,…

Robotics · Computer Science 2026-03-20 Swagat Padhan , Lakshya Jain , Bhavya Minesh Shah , Omkar Patil , Thao Nguyen , Nakul Gopalan

Exploiting visual groundings for language understanding has recently been drawing much attention. In this work, we study visually grounded grammar induction and learn a constituency parser from both unlabeled text and its visual groundings.…

Computation and Language · Computer Science 2020-12-08 Yanpeng Zhao , Ivan Titov

Localizing the exact pathological regions in a given medical scan is an important imaging problem that traditionally requires a large amount of bounding box ground truth annotations to be accurately solved. However, there exist alternative,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-31 Konstantinos Vilouras , Pedro Sanchez , Alison Q. O'Neil , Sotirios A. Tsaftaris

In clinical radiology reports, doctors capture important information about the patient's health status. They convey their observations from raw medical imaging data about the inner structures of a patient. As such, formulating reports…

Computer Vision and Pattern Recognition · Computer Science 2022-10-10 Constantin Seibold , Simon Reiß , Saquib Sarfraz , Matthias A. Fink , Victoria Mayer , Jan Sellner , Moon Sung Kim , Klaus H. Maier-Hein , Jens Kleesiek , Rainer Stiefelhagen

Vision--language models (VLMs) for radiology report generation (RRG) can produce long-form chest CT reports from volumetric scans and show strong potential to improve radiology workflow efficiency and consistency. However, existing methods…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Chenyu Wang , Weicheng Dai , Han Liu , Wenchao Li , Kayhan Batmanghelich

Automatic radiology report generation has been an attracting research problem towards computer-aided diagnosis to alleviate the workload of doctors in recent years. Deep learning techniques for natural image captioning are successfully…

Computer Vision and Pattern Recognition · Computer Science 2020-02-20 Yixiao Zhang , Xiaosong Wang , Ziyue Xu , Qihang Yu , Alan Yuille , Daguang Xu

Despite significant advancements in adapting Large Language Models (LLMs) for radiology report generation (RRG), clinical adoption remains challenging due to difficulties in accurately mapping pathological and anatomical features to their…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Qilong Xing , Zikai Song , Youjia Zhang , Na Feng , Junqing Yu , Wei Yang

Radiology Report Generation (RRG) aims to automatically generate diagnostic reports from radiology images. To achieve this, existing methods have leveraged the powerful cross-modal generation capabilities of Multimodal Large Language Models…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Jiechao Gao , Chang Liu , Yuangang Li

We address the problem of phrase grounding by lear ing a multi-level common semantic space shared by the textual and visual modalities. We exploit multiple levels of feature maps of a Deep Convolutional Neural Network, as well as…

Computer Vision and Pattern Recognition · Computer Science 2019-05-31 Hassan Akbari , Svebor Karaman , Surabhi Bhargava , Brian Chen , Carl Vondrick , Shih-Fu Chang

Radiology report generation is critical for efficiency but current models lack the structured reasoning of experts, hindering clinical trust and explainability by failing to link visual findings to precise anatomical locations. This paper…

Artificial Intelligence · Computer Science 2026-03-03 Peiyuan Jing , Kinhei Lee , Zhenxuan Zhang , Huichi Zhou , Zhengqing Yuan , Zhifan Gao , Lei Zhu , Giorgos Papanastasiou , Yingying Fang , Guang Yang

Automated radiology report generation has gained increasing attention with the rise of deep learning and large language models. However, fully generative approaches often suffer from hallucinations and lack clinical grounding, limiting…

Quantitative Methods · Quantitative Biology 2026-05-01 Himadri S Samanta

Multi-modal large language models have demonstrated impressive performance across various tasks in different modalities. However, existing multi-modal models primarily emphasize capturing global information within each modality while…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Zhaowei Li , Qi Xu , Dong Zhang , Hang Song , Yiqing Cai , Qi Qi , Ran Zhou , Junting Pan , Zefeng Li , Van Tu Vu , Zhida Huang , Tao Wang