English
Related papers

Related papers: Enhanced Contrastive Learning with Multi-view Long…

200 papers

Cross-modal medical image-report retrieval task plays a significant role in clinical diagnosis and various medical generative tasks. Eliminating heterogeneity between different modalities to enhance semantic consistency is the key challenge…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Zeqiang Wei , Kai Jin , Xiuzhuang Zhou

Self-supervised learning has proven to be an effective way to learn representations in domains where annotated labels are scarce, such as medical imaging. A widely adopted framework for this purpose is contrastive learning and it has been…

Computer Vision and Pattern Recognition · Computer Science 2024-02-23 Hugo Figueiras , Helena Aidos , Nuno Cruz Garcia

Recent advancements in self-supervised learning have demonstrated that effective visual representations can be learned from unlabeled images. This has led to increased interest in applying self-supervised learning to the medical domain,…

Computer Vision and Pattern Recognition · Computer Science 2023-04-10 Xiangyi Yan , Junayed Naushad , Chenyu You , Hao Tang , Shanlin Sun , Kun Han , Haoyu Ma , James Duncan , Xiaohui Xie

Multi-modal magnetic resonance imaging (MRI) is essential for providing complementary information about brain anatomy and pathology, leading to more accurate diagnoses. However, obtaining high-quality multi-modal MRI in a clinical setting…

Image and Video Processing · Electrical Eng. & Systems 2025-04-15 Minjoo Lim , Bogyeong Kang , Tae-Eui Kam

The deep learning technique has been shown to be effectively addressed several image analysis tasks in the computer-aided diagnosis scheme for mammography. The training of an efficacious deep learning model requires large data with diverse…

Computer Vision and Pattern Recognition · Computer Science 2023-09-08 Zheren Li , Zhiming Cui , Lichi Zhang , Sheng Wang , Chenjin Lei , Xi Ouyang , Dongdong Chen , Xiangyu Zhao , Yajia Gu , Zaiyi Liu , Chunling Liu , Dinggang Shen , Jie-Zhi Cheng

Large vision-language models (VLMs) have evolved from general-purpose applications to specialized use cases such as in the clinical domain, demonstrating potential for decision support in radiology. One promising application is assisting…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Zhifan Jiang , Dong Yang , Vishwesh Nath , Abhijeet Parida , Nishad P. Kulkarni , Ziyue Xu , Daguang Xu , Syed Muhammad Anwar , Holger R. Roth , Marius George Linguraru

Radiology report generation (RRG) requires advanced medical image analysis, effective temporal reasoning, and accurate text generation. While multimodal large language models (MLLMs) align with pre-trained vision encoders to enhance…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Xi Zhang , Zaiqiao Meng , Jake Lever , Edmond S. L. Ho

Radiology report generation (RRG) aims to automatically generate free-text descriptions from clinical radiographs, e.g., chest X-Ray images. RRG plays an essential role in promoting clinical automation and presents significant help to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Chang Liu , Yuanhe Tian , Yan Song

Analysis of cardiac ultrasound images is commonly performed in routine clinical practice for quantification of cardiac function. Its increasing automation frequently employs deep learning networks that are trained to predict disease or…

Computer Vision and Pattern Recognition · Computer Science 2021-08-09 Agisilaos Chartsias , Shan Gao , Angela Mumith , Jorge Oliveira , Kanwal Bhatia , Bernhard Kainz , Arian Beqiri

Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high dependence on large-scale, high-quality paired data and the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-16 Zehan Wang , Ziang Zhang , Luping Liu , Yang Zhao , Haifeng Huang , Tao Jin , Zhou Zhao

In clinics, a radiology report is crucial for guiding a patient's treatment. However, writing radiology reports is a heavy burden for radiologists. To this end, we present an automatic, multi-modal approach for report generation from a…

Image and Video Processing · Electrical Eng. & Systems 2022-06-02 Shuxin Yang , Xian Wu , Shen Ge , S. Kevin Zhou , Li Xiao

The integration of different imaging modalities, such as structural, diffusion tensor, and functional magnetic resonance imaging, with deep learning models has yielded promising outcomes in discerning phenotypic characteristics and…

Image and Video Processing · Electrical Eng. & Systems 2024-10-08 Zhiyuan Li , Hailong Li , Anca L. Ralescu , Jonathan R. Dillman , Mekibib Altaye , Kim M. Cecil , Nehal A. Parikh , Lili He

Electronic health record (EHR) systems contain a wealth of multimodal clinical data including structured data like clinical codes and unstructured data such as clinical notes. However, many existing EHR-focused studies has traditionally…

Machine Learning · Statistics 2025-08-20 Tianxi Cai , Feiqing Huang , Ryumei Nakada , Linjun Zhang , Doudou Zhou

This study investigates the integration of diverse patient data sources into multimodal language models for automated chest X-ray (CXR) report generation. Traditionally, CXR report generation relies solely on CXR images and limited…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Aaron Nicolson , Shengyao Zhuang , Jason Dowling , Bevan Koopman

The vision-language modeling capability of multi-modal large language models has attracted wide attention from the community. However, in medical domain, radiology report generation using vision-language models still faces significant…

Computer Vision and Pattern Recognition · Computer Science 2024-08-23 Yuhao Wang , Chao Hao , Yawen Cui , Xinqi Su , Weicheng Xie , Tao Tan , Zitong Yu

The rapid advancements in large language models (LLMs) have unlocked their potential for multimodal tasks, where text and visual data are processed jointly. However, applying LLMs to medical imaging, particularly for chest X-rays (CXR),…

Image and Video Processing · Electrical Eng. & Systems 2025-02-11 Nicholas Evans , Stephen Baker , Miles Reed

Representation learning offers a conduit to elucidate distinctive features within the latent space and interpret the deep models. However, the randomness of lesion distribution and the complexity of low-quality factors in medical images…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Qingshan Hou , Shuai Cheng , Peng Cao , Jinzhu Yang , Xiaoli Liu , Osmar R. Zaiane , Yih Chung Tham

In breast cancer detection and diagnosis, the longitudinal analysis of mammogram images is crucial. Contemporary models excel in detecting temporal imaging feature changes, thus enhancing the learning process over sequential imaging exams.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Zhengbo Zhou , Degan Hao , Dooman Arefan , Margarita Zuley , Jules Sumkin , Shandong Wu

Multi-view representation learning has developed rapidly over the past decades and has been applied in many fields. However, most previous works assumed that each view is complete and aligned. This leads to an inevitable deterioration in…

Computer Vision and Pattern Recognition · Computer Science 2022-11-10 Yiming Wang , Dongxia Chang , Zhiqiang Fu , Jie Wen , Yao Zhao

Automated radiology report generation (RRG) holds potential to reduce the workload of radiologists, and recent advances in multimodal large language models (MLLMs) have enabled multimodal chest X-ray (CXR) report generation. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Jonggwon Park , Byungmu Yoon , Soobum Kim , Kyoyun Choi