English
Related papers

Related papers: MedVersa: A Generalist Foundation Model for Medica…

200 papers

Multi-modal representation methods have achieved advanced performance in medical applications by extracting more robust features from multi-domain data. However, existing methods usually need to train additional branches for downstream…

Computer Vision and Pattern Recognition · Computer Science 2023-03-16 Weijian Huang , Hao Yang , Cheng Li , Mingtong Dai , Rui Yang , Shanshan Wang

We introduce MultiMedEval, an open-source toolkit for fair and reproducible evaluation of large, medical vision-language models (VLM). MultiMedEval comprehensively assesses the models' performance on a broad array of six multi-modal tasks,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Corentin Royer , Bjoern Menze , Anjany Sekuboyina

Recent advancements in Large Multimodal Models (LMMs) have attracted interest in their generalization capability with only a few samples in the prompt. This progress is particularly relevant to the medical domain, where the quality and…

Computation and Language · Computer Science 2024-05-06 Seonhee Cho , Choonghan Kim , Jiho Lee , Chetan Chilkunda , Sujin Choi , Joo Heung Yoon

The integration of AI-assisted biomedical image analysis into clinical practice demands AI-generated findings that are not only accurate but also interpretable to clinicians. However, existing biomedical AI models generally lack the ability…

Automatic radiology report generation can alleviate the workload for physicians and minimize regional disparities in medical resources, therefore becoming an important topic in the medical image analysis field. It is a challenging task, as…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Xinyi Wang , Grazziela Figueredo , Ruizhe Li , Wei Emma Zhang , Weitong Chen , Xin Chen

Medical image recognition serves as a key way to aid in clinical diagnosis, enabling more accurate and timely identification of diseases and abnormalities. Vision transformer-based approaches have proven effective in handling various…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Zunhui Xia , Hongxing Li , Libin Lan

Biomedical multimodal assistants have the potential to unify radiology, pathology, and clinical-text reasoning, yet a critical deployment gap remains: top-performing systems are either closed-source or computationally prohibitive,…

Computation and Language · Computer Science 2026-03-03 Kai Zhang , Zhengqing Yuan , Cheng Peng , Songlin Zhao , Mengxian Lyu , Ziyi Chen , Yanfang Ye , Wei Liu , Ying Zhang , Kaleb E Smith , Lifang He , Lichao Sun , Yonghui Wu

Recent advancements in mixed-modal generative have opened new avenues for developing unified biomedical assistants capable of analyzing biomedical images, answering complex questions about them, and generating multimodal patient reports.…

Artificial Intelligence · Computer Science 2025-04-24 Hritik Bansal , Daniel Israel , Siyan Zhao , Shufan Li , Tung Nguyen , Aditya Grover

The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of LLMs. Unlike earlier specialists, existing MLLMs are evolving towards a Multimodal Generalist paradigm. Initially…

Biomedical data is inherently multimodal, comprising physical measurements and natural language narratives. A generalist biomedical AI model needs to simultaneously process different modalities of data, including text and images. Therefore,…

Pathology images are crucial for diagnosing and managing various diseases by visualizing cellular and tissue-level abnormalities. Recent advancements in artificial intelligence (AI), particularly multimodal models like ChatGPT, have shown…

Human-Computer Interaction · Computer Science 2024-09-25 Mianxin Liu , Jianfeng Wu , Fang Yan , Hongjun Li , Wei Wang , Shaoting Zhang , Zhe Wang

Multimodal models are expected to be a critical component to future advances in artificial intelligence. This field is starting to grow rapidly with a surge of new design elements motivated by the success of foundation models in natural…

Computation and Language · Computer Science 2024-06-11 Sai Munikoti , Ian Stewart , Sameera Horawalavithana , Henry Kvinge , Tegan Emerson , Sandra E Thompson , Karl Pazdernik

Current artificial intelligence models for medical imaging are predominantly single modality and single disease. Attempts to create multimodal and multi-disease models have resulted in inconsistent clinical accuracy. Furthermore, training…

With machine learning successfully applied to new daunting problems almost every day, general AI starts looking like an attainable goal. However, most current research focuses instead on important but narrow applications, such as image…

Machine Learning · Computer Science 2017-03-29 Marco Baroni , Armand Joulin , Allan Jabri , Germàn Kruszewski , Angeliki Lazaridou , Klemen Simonic , Tomas Mikolov

The discussions around Artificial Intelligence (AI) and medical imaging are centered around the success of deep learning algorithms. As new algorithms enter the market, it is important for practicing radiologists to understand the pitfalls…

Image and Video Processing · Electrical Eng. & Systems 2022-11-28 Rishi Gadepally , Andrew Gomella , Eric Gingold , Paras Lakhani

Accurate medical diagnosis often involves progressive visual focusing and iterative reasoning, characteristics commonly observed in clinical workflows. While recent vision-language models demonstrate promising chain-of-thought (CoT)…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Chunzheng Zhu , Yangfang Lin , Shen Chen , Yijun Wang , Jianxin Lin

Deep learning algorithms have demonstrated remarkable efficacy in various medical image analysis (MedIA) applications. However, recent research highlights a performance disparity in these algorithms when applied to specific subgroups, such…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Zikang Xu , Jun Li , Qingsong Yao , Han Li , Mingyue Zhao , S. Kevin Zhou

Medical data poses a daunting challenge for AI algorithms: it exists in many different modalities, experiences frequent distribution shifts, and suffers from a scarcity of examples and labels. Recent advances, including transformers and…

Foundation models, often pre-trained with large-scale data, have achieved paramount success in jump-starting various vision and language applications. Recent advances further enable adapting foundation models in downstream tasks efficiently…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Dequan Wang , Xiaosong Wang , Lilong Wang , Mengzhang Li , Qian Da , Xiaoqiang Liu , Xiangyu Gao , Jun Shen , Junjun He , Tian Shen , Qi Duan , Jie Zhao , Kang Li , Yu Qiao , Shaoting Zhang

Medical Visual Language Models have shown great potential in various healthcare applications, including medical image captioning and diagnostic assistance. However, most existing models rely on text-based instructions, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Tan-Hanh Pham , Chris Ngo , Trong-Duong Bui , Minh Luu Quang , Tan-Huong Pham , Truong-Son Hy