English
Related papers

Related papers: DeViDe: Faceted medical knowledge for improved med…

200 papers

Radiology report generation aims to automatically generate a clinically accurate and coherent paragraph from the X-ray image, which could relieve radiologists from the heavy burden of report writing. Although various image caption methods…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Zhongzhen Huang , Xiaofan Zhang , Shaoting Zhang

Understanding 3D medical image volumes is a critical task in the medical domain. However, existing 3D convolution and transformer-based methods have limited semantic understanding of an image volume and also need a large set of volumes for…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Qiuhui Chen , Huping Ye , Yi Hong

Vision-and-language models (VLMs) have been increasingly explored in the medical domain, particularly following the success of CLIP in general domain. However, unlike the relatively straightforward pairing of 2D images and text, curating…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Ziyang Zhang , Yang Yu , Xulei Yang , Si Yong Yeo

Interpretability and small labelled datasets are key issues in the practical application of deep learning, particularly in areas such as medicine. In this paper, we present a semi-supervised technique that addresses both these issues by…

Computer Vision and Pattern Recognition · Computer Science 2018-04-13 Jarrel Seah , Jennifer Tang , Andy Kitchen , Jonathan Seah

Medical image visual question answering (VQA) is a task to answer clinical questions, given a radiographic image, which is a challenging problem that requires a model to integrate both vision and language information. To solve medical VQA…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Pengfei Li , Gang Liu , Lin Tan , Jinying Liao , Shenjun Zhong

Medical Visual Language Models have shown great potential in various healthcare applications, including medical image captioning and diagnostic assistance. However, most existing models rely on text-based instructions, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Tan-Hanh Pham , Chris Ngo , Trong-Duong Bui , Minh Luu Quang , Tan-Huong Pham , Truong-Son Hy

Medical image analysis is crucial in modern radiological diagnostics, especially given the exponential growth in medical imaging data. The demand for automated report generation systems has become increasingly urgent. While prior research…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Hao Chen , Wei Zhao , Yingli Li , Tianyang Zhong , Yisong Wang , Youlan Shang , Lei Guo , Junwei Han , Tianming Liu , Jun Liu , Tuo Zhang

Vision-language pre-training (VLP) has great potential for developing multifunctional and general medical diagnostic capabilities. However, aligning medical images with a low signal-to-noise ratio (SNR) to reports with a high SNR presents a…

Image and Video Processing · Electrical Eng. & Systems 2025-08-07 Weiwei Cao , Jianpeng Zhang , Zhongyi Shui , Sinuo Wang , Zeli Chen , Xi Li , Le Lu , Xianghua Ye , Tingbo Liang , Qi Zhang , Ling Zhang

Referring image segmentation is a fundamental vision-language task that aims to segment out an object referred to by a natural language expression from an image. One of the key challenges behind this task is leveraging the referring…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Zhao Yang , Jiaqi Wang , Yansong Tang , Kai Chen , Hengshuang Zhao , Philip H. S. Torr

Recent advancements in large-scale Vision Transformers have made significant strides in improving pre-trained models for medical image segmentation. However, these methods face a notable challenge in acquiring a substantial amount of…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Yiqing Wang , Zihan Li , Jieru Mei , Zihao Wei , Li Liu , Chen Wang , Shengtian Sang , Alan Yuille , Cihang Xie , Yuyin Zhou

Multi-modal medical imaging enables comprehensive diagnostics, yet current foundation models process 2D (e.g. X-ray) and 3D (e.g. CT) data with separate, dimensionality-specific architectures. We present MultiMedVision, a unified framework…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Frank Li , Bardia Khosravi , Mohammadreza Chavoshi , Young Seok Jeon , Theo Dapamede , Hari Trivedi , Janice Newsome , Judy Gichoya

Vision-language models (VLMs) show strong potential for complex diagnostic tasks in medical imaging. However, applying VLMs to multi-organ medical imaging introduces two principal challenges: (1) modality-specific vision-language alignment…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Haowen Zhu , Ning Yin , Xiaogen Zhou

Chest X-ray report generation aims to reduce radiologists' workload by automatically producing high-quality preliminary reports. A critical yet underexplored aspect of this task is the effective use of patient-specific prior knowledge --…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Kang Liu , Zhuoqi Ma , Zikang Fang , Yunan Li , Kun Xie , Qiguang Miao

Mammography is the primary imaging tool for breast cancer diagnosis. Despite significant strides in applying deep learning to interpret mammography images, efforts that focus predominantly on visual features often struggle with…

Image and Video Processing · Electrical Eng. & Systems 2024-09-25 Xin Wei , Yaling Tao , Changde Du , Gangming Zhao , Yizhou Yu , Jinpeng Li

Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable promise in generating visually grounded responses. However, their application in the medical domain is hindered by unique challenges. For instance, most VLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Lingxiao Luo , Bingda Tang , Xuanzhong Chen , Rong Han , Ting Chen

Medical imaging is an invaluable resource in medicine as it enables to peer inside the human body and provides scientists and physicians with a wealth of information indispensable for understanding, modelling, diagnosis, and treatment of…

Image and Video Processing · Electrical Eng. & Systems 2022-08-25 Hanene Ben Yedder , Ben Cardoen , Ghassan Hamarneh

Localized image captioning has made significant progress with models like the Describe Anything Model (DAM), which can generate detailed region-specific descriptions without explicit region-text supervision. However, such capabilities have…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Xi Xiao , Yunbei Zhang , Thanh-Huy Nguyen , Ba-Thinh Lam , Janet Wang , Lin Zhao , Jihun Hamm , Tianyang Wang , Xingjian Li , Xiao Wang , Hao Xu , Tianming Liu , Min Xu

Background: Deep learning has significantly advanced medical image analysis, with Vision Transformers (ViTs) offering a powerful alternative to convolutional models by modeling long-range dependencies through self-attention. However, ViTs…

The inability to interpret the model prediction in semantically and visually meaningful ways is a well-known shortcoming of most existing computer-aided diagnosis methods. In this paper, we propose MDNet to establish a direct multimodal…

Computer Vision and Pattern Recognition · Computer Science 2017-07-11 Zizhao Zhang , Yuanpu Xie , Fuyong Xing , Mason McGough , Lin Yang

Due to numerous hardware shortcomings, medical image acquisition devices are susceptible to producing low-quality (i.e., low contrast, inappropriate brightness, noisy, etc.) images. Regrettably, perceptually degraded images directly impact…

Image and Video Processing · Electrical Eng. & Systems 2025-03-12 S M A Sharif , Rizwan Ali Naqvi , Mithun Biswas , Woong-Kee Loh