English
Related papers

Related papers: MEDBind: Unifying Language and Multimodal Medical …

200 papers

Multimodal representation learning has demonstrated remarkable potential in enabling models to process and integrate diverse data modalities, such as text and images, for improved understanding and performance. While the medical domain can…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Shuvendu Roy , Franklin Ogidi , Ali Etemad , Elham Dolatabadi , Arash Afkanpour

Vision-language models (VLMs) show strong potential for complex diagnostic tasks in medical imaging. However, applying VLMs to multi-organ medical imaging introduces two principal challenges: (1) modality-specific vision-language alignment…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Haowen Zhu , Ning Yin , Xiaogen Zhou

Recent advances in multimodal large language models (MLLMs) have significantly improved medical AI, enabling it to unify the understanding of visual and textual information. However, as medical knowledge continues to evolve, it is critical…

Artificial Intelligence · Computer Science 2025-08-08 Dexuan Xu , Jieyi Wang , Zhongyan Chai , Yongzhi Cao , Hanpin Wang , Huamin Zhang , Yu Huang

Recent advances in large language models (LLMs) have enabled the development of multimodal medical AI. While models such as MedGemini achieve high accuracy on VQA tasks like USMLE MM, their performance on ECG based tasks remains limited,…

Machine Learning · Computer Science 2026-02-12 Junichiro Takahashi , Masataka Sato , Satoshi Kodeta , Norihiko Takeda

The imperative for early detection of type 2 diabetes mellitus (T2DM) is challenged by its asymptomatic onset and dependence on suboptimal clinical diagnostic tests, contributing to its widespread global prevalence. While research into…

Machine Learning · Computer Science 2024-12-17 Sanjana Gundapaneni , Zhuo Zhi , Miguel Rodrigues

Accurate detection and localization of traumatic injuries in abdominal CT scans remains a critical challenge in emergency radiology, primarily due to severe scarcity of annotated medical data. This paper presents a label-efficient approach…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Shivam Chaudhary , Sheethal Bhat , Andreas Maier

Neuroimaging techniques including functional magnetic resonance imaging (fMRI) and electroencephalogram (EEG) have shown promise in detecting functional abnormalities in various brain disorders. However, existing studies often focus on a…

Image and Video Processing · Electrical Eng. & Systems 2024-10-01 Xinxu Wei , Kanhao Zhao , Yong Jiao , Nancy B. Carlisle , Hua Xie , Gregory A. Fonzo , Yu Zhang

Biomedical word embeddings are usually pre-trained on free text corpora with neural methods that capture local and global distributional properties. They are leveraged in downstream tasks using various neural architectures that are designed…

Computation and Language · Computer Science 2021-07-26 Jiho Noh , Ramakanth Kavuluru

Multi-modal models require aligned, shared embedding spaces. However, common CLIP-based approaches need large amounts of samples and do not natively support 3D or tabular data, both of which are crucial in the medical domain. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Jakob Krogh Petersen , Valdemar Licht , Mads Nielsen , Asbjørn Munk

This study investigates the integration of diverse patient data sources into multimodal language models for automated chest X-ray (CXR) report generation. Traditionally, CXR report generation relies solely on CXR images and limited…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Aaron Nicolson , Shengyao Zhuang , Jason Dowling , Bevan Koopman

Medicine is inherently multimodal and multitask, with diverse data modalities spanning text, imaging. However, most models in medical field are unimodal single tasks and lack good generalizability and explainability. In this study, we…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Lijian Xu , Hao Sun , Ziyu Ni , Hongsheng Li , Shaoting Zhang

Deep learning models exhibit state-of-the-art performance for many predictive healthcare tasks using electronic health records (EHR) data, but these models typically require training data volume that exceeds the capacity of most healthcare…

Machine Learning · Computer Science 2018-10-24 Edward Choi , Cao Xiao , Walter F. Stewart , Jimeng Sun

We propose Medical Entity Definition-based Sentence Embedding (MED-SE), a novel unsupervised contrastive learning framework designed for clinical texts, which exploits the definitions of medical entities. To this end, we conduct an…

Machine Learning · Computer Science 2022-12-12 Hyeonbin Hwang , Haanju Yoo , Yera Choi

In recent years, cross-modal retrieval using images and text has become an active area of research, especially in the medical domain. The abundance of data in various modalities in this field has led to a growing importance of cross-modal…

Information Retrieval · Computer Science 2025-12-09 Jaewon Ahn , Woosung Jang , Beakcheol Jang

Conventional machine learning models, particularly tree-based approaches, have demonstrated promising performance across various clinical prediction tasks using electronic health record (EHR) data. Despite their strengths, these models…

Computation and Language · Computer Science 2025-05-26 Sara Ketabi , Dhanesh Ramachandram

Language models can be viewed as functions that embed text into Euclidean space, where the quality of the embedding vectors directly determines model performance, training such neural networks involves various uncertainties. This paper…

Computation and Language · Computer Science 2025-03-31 Yifei Duan , Raphael Shang , Deng Liang , Yongqiang Cai

Despite recent progress in text-prompt-based medical image segmentation, these methods are limited to single-round dialogues and fail to support multi-round reasoning, which is important for medical education scenarios. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Qinyue Tong , Ziqian Lu , Jun Liu , Rui Zuo , Zheming Lu , Yueming Jin

Depression is a widespread mental health disorder, yet its automatic detection remains challenging. Prior work has explored unimodal and multimodal approaches, with multimodal systems showing promise by leveraging complementary signals.…

Artificial Intelligence · Computer Science 2026-03-24 Annisaa Fitri Nurfidausi , Eleonora Mancini , Paolo Torroni

The rapid advancement of Multimodal Large Language Models (MLLMs) has extended CLIP-based frameworks to produce powerful, universal embeddings for retrieval tasks. However, existing methods primarily focus on natural images, offering…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Weijian Jian , Yajun Zhang , Dawei Liang , Chunyu Xie , Yixiao He , Dawei Leng , Yuhui Yin

Multimodal models have achieved remarkable success in natural image segmentation, yet they often underperform when applied to the medical domain. Through extensive study, we attribute this performance gap to the challenges of multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Wenjun Yu , Yinchen Zhou , Jia-Xuan Jiang , Shubin Zeng , Yuee Li , Zhong Wang