English
Related papers

Related papers: M3T: Multi-Modal Medical Transformer to bridge Cli…

200 papers

Accelerated multi-modal magnetic resonance (MR) imaging is a new and effective solution for fast MR imaging, providing superior performance in restoring the target modality from its undersampled counterpart with guidance from an auxiliary…

Image and Video Processing · Electrical Eng. & Systems 2022-05-12 Chun-Mei Feng , Yunlu Yan , Geng Chen , Yong Xu , Ling Shao , Huazhu Fu

When it comes to clinical images, automatic segmentation has a wide variety of applications and a considerable diversity of input domains, such as different types of Magnetic Resonance Images (MRIs) and Computerized Tomography (CT) scans.…

Image and Video Processing · Electrical Eng. & Systems 2024-02-28 Matteo Bastico , David Ryckelynck , Laurent Corté , Yannick Tillier , Etienne Decencière

Automated medical report generation has become increasingly important in medical analysis. It can produce computer-aided diagnosis descriptions and thus significantly alleviate the doctors' work. Inspired by the huge success of neural…

Computer Vision and Pattern Recognition · Computer Science 2023-08-11 Keqiang Fan , Xiaohao Cai , Mahesan Niranjan

The capability to jointly process multi-modal information is becoming an essential task. However, the limited number of paired multi-modal data and the large computational requirements in multi-modal learning hinder the development. We…

Computation and Language · Computer Science 2025-06-09 Minsu Kim , Jee-weon Jung , Hyeongseop Rha , Soumi Maiti , Siddhant Arora , Xuankai Chang , Shinji Watanabe , Yong Man Ro

Clinical diagnosis is a highly specialized discipline requiring both domain expertise and strict adherence to rigorous guidelines. While current AI-driven medical research predominantly focuses on knowledge graphs or natural text…

Machine Learning · Computer Science 2025-12-12 Haolin Li , Tianjie Dai , Zhe Chen , Siyuan Du , Jiangchao Yao , Ya Zhang , Yanfeng Wang

The inability to interpret the model prediction in semantically and visually meaningful ways is a well-known shortcoming of most existing computer-aided diagnosis methods. In this paper, we propose MDNet to establish a direct multimodal…

Computer Vision and Pattern Recognition · Computer Science 2017-07-11 Zizhao Zhang , Yuanpu Xie , Fuyong Xing , Mason McGough , Lin Yang

X-ray medical report generation is one of the important applications of artificial intelligence in healthcare. With the support of large foundation models, the quality of medical report generation has significantly improved. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Futian Wang , Yuhan Qiao , Xiao Wang , Fuling Wang , Yuxiang Zhang , Dengdi Sun

Cross-modal medical image translation is an essential task for synthesizing missing modality data for clinical diagnosis. However, current learning-based techniques have limitations in capturing cross-modal and global features, restricting…

Computer Vision and Pattern Recognition · Computer Science 2023-10-05 Xuhang Chen , Chi-Man Pun , Shuqiang Wang

3D medical image analysis is essential for modern healthcare, yet traditional task-specific models are inadequate due to limited generalizability across diverse clinical scenarios. Multimodal large language models (MLLMs) offer a promising…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Yiming Shi , Xun Zhu , Kaiwen Wang , Ying Hu , Chenyi Guo , Miao Li , Ji Wu

While deep learning methods have shown great success in medical image analysis, they require a number of medical images to train. Due to data privacy concerns and unavailability of medical annotators, it is oftentimes very difficult to…

Image and Video Processing · Electrical Eng. & Systems 2020-10-08 Yue Yang , Pengtao Xie

Three-dimensional medical image data and computer-aided decision making, particularly using deep learning, are becoming increasingly important in the medical field. To aid in these developments we introduce PR3DICTR: Platform for Research…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Daniel C. MacRae , Luuk van der Hoek , Robert van der Wal , Suzanne P. M. de Vette , Hendrike Neh , Baoqiang Ma , Peter M. A. van Ooijen , Lisanne V. van Dijk

In this paper, we propose a framework that incorporates experts diagnostics and insights into the analysis of Optical Coherence Tomography (OCT) using multi-modal learning. To demonstrate the effectiveness of this approach, we create a…

Image and Video Processing · Electrical Eng. & Systems 2022-03-22 Y. Logan , K. Kokilepersaud , G. Kwon , G. AlRegib , C. Wykoff , H. Yu

We present a cross-modality generation framework that learns to generate translated modalities from given modalities in MR images without real acquisition. Our proposed method performs NeuroImage-to-NeuroImage translation (abbreviated as…

Computer Vision and Pattern Recognition · Computer Science 2018-09-12 Qianye Yang , Nannan Li , Zixu Zhao , Xingyu Fan , Eric I-Chao Chang , Yan Xu

Radiologic diagnostic errors-under-reading errors, inattentional blindness, and communication failures-remain prevalent in clinical practice. These issues often stem from missed localized abnormalities, limited global context, and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Yuheng Li , Yenho Chen , Yuxiang Lai , Jike Zhong , Vanessa Wildman , Xiaofeng Yang

Accurate segmentation of organs and tumors in CT and MRI scans is essential for diagnosis, treatment planning, and disease monitoring. While deep learning has advanced automated segmentation, most models remain task-specific, lacking…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Yuheng Li , Yizhou Wu , Yuxiang Lai , Mingzhe Hu , Xiaofeng Yang

While multimodal data integrating diverse imaging and clinical tabular records is crucial for accurate medical diagnosis, the arbitrary absence of specific modalities is prevalent in clinical practice, severely degrading the performance of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Tianling Liu , Lequan Yu , Tong Han , Liang Wan

Images in the medical domain are fundamentally different from the general domain images. Consequently, it is infeasible to directly employ general domain Visual Question Answering (VQA) models for the medical domain. Additionally, medical…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Yash Khare , Viraj Bagal , Minesh Mathew , Adithi Devi , U Deva Priyakumar , CV Jawahar

In this paper, an innovative multi-modal deep learning model is proposed to deeply integrate heterogeneous information from medical images and clinical reports. First, for medical images, convolutional neural networks were used to extract…

Machine Learning · Computer Science 2024-05-29 Ziyan Yao , Fei Lin , Sheng Chai , Weijie He , Lu Dai , Xinghui Fei

Multi-modal large language models (MLLMs) have shown promise in advancing healthcare. However, most existing models remain confined to single-image understanding, which greatly limits their applicability in clinical workflows. In practice,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Zhen Chen , Yihang Fu , Gabriel Madera , Mauro Giuffre , Serina Applebaum , Hyunjae Kim , Hua Xu , Qingyu Chen

Multimodal machine translation (MMT) aims to improve translation quality by incorporating information from other modalities, such as vision. Previous MMT systems mainly focus on better access and use of visual information and tend to…

Computation and Language · Computer Science 2023-09-06 Yaoming Zhu , Zewei Sun , Shanbo Cheng , Luyang Huang , Liwei Wu , Mingxuan Wang