English
Related papers

Related papers: MedVersa: A Generalist Foundation Model for Medica…

200 papers

With the advent of Vision-Language Models (VLMs), medical artificial intelligence (AI) has experienced significant technological progress and paradigm shifts. This survey provides an extensive review of recent advancements in Medical…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Beria Chingnabe Kalpelbe , Angel Gabriel Adaambiik , Wei Peng

Artificial Intelligence (AI) has become commonplace to solve routine everyday tasks. Because of the exponential growth in medical imaging data volume and complexity, the workload on radiologists is steadily increasing. We project that the…

Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.g., "Is this normal or abnormal?") or qualitative descriptive tasks. However, clinical decision-making often relies on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Yongcheng Yao , Yongshuo Zong , Raman Dutt , Yongxin Yang , Sotirios A Tsaftaris , Timothy Hospedales

Large language models (LLMs) have demonstrated strong performance and rapid progress in a wide range of medical reasoning tasks. However, their sequential autoregressive decoding forces inherently parallel clinical reasoning, such as…

Machine Learning · Computer Science 2026-04-17 Jianwen Chen , Xinyu Yang , Peng Xia , Arian Azarang , Yueh Z Lee , Gang Li , Hongtu Zhu , Yun Li , Beidi Chen , Huaxiu Yao

In the past decade, deep learning (DL) has achieved unprecedented success in numerous fields including computer vision, natural language processing, and healthcare. In particular, DL is experiencing an increasing development in applications…

Computer Vision and Pattern Recognition · Computer Science 2020-07-15 Liang Zhang , Johann Li , Ping Li , Xiaoyuan Lu , Peiyi Shen , Guangming Zhu , Syed Afaq Shah , Mohammed Bennarmoun , Kun Qian , Björn W. Schuller

Medical images often contain multiple labels with imbalanced distributions and co-occurrence, leading to bias in multi-label medical image classification. Close collaboration between medical professionals and machine learning practitioners…

Human-Computer Interaction · Computer Science 2025-07-30 Shaohan Shi , Yuheng Shao , Haoran Jiang , Yunjie Yao , Zhijun Zhang , Xu Ding , Quan Li

Medical image analysis is essential in modern healthcare. Deep learning has redirected research focus toward complex medical multimodal tasks, including report generation and visual question answering. Traditional task-specific models often…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Yiming Shi , Shaoshuai Yang , Xun Zhu , Haoyu Wang , Xiangling Fu , Miao Li , Ji Wu

Medical artificial general intelligence (MAGI) enables one foundation model to solve different medical tasks, which is very practical in the medical domain. It can significantly reduce the requirement of large amounts of task-specific data…

Research in medical imaging primarily focuses on discrete data representations that poorly scale with grid resolution and fail to capture the often continuous nature of the underlying signal. Neural Fields (NFs) offer a powerful alternative…

Image and Video Processing · Electrical Eng. & Systems 2026-03-06 Paul Friedrich , Florentin Bieder , Julian McGinnis , Julia Wolleb , Daniel Rueckert , Philippe C. Cattin

Generalist multimodal large language models (MLLMs) have achieved impressive performance across a wide range of vision-language tasks. However, their performance on medical tasks, particularly in zero-shot settings where generalization is…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Guimeng Liu , Tianze Yu , Somayeh Ebrahimkhani , Lin Zhi Zheng Shawn , Kok Pin Ng , Ngai-Man Cheung

Diagnosing and managing oral diseases necessitate advanced visual interpretation across diverse imaging modalities and integrated information synthesis. While current AI models excel at isolated tasks, they often fall short in addressing…

Large language models perform well on many medical QA benchmarks, but real clinical reasoning often requires integrating evidence across multiple images rather than interpreting a single view. We introduce MedThinkVQA, an expert-annotated…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Zonghai Yao , Benlu Wang , Yifan Zhang , Junda Wang , Iris Xia , Zhipeng Tang , Shuo Han , Feiyun Ouyang , Zhichao Yang , Arman Cohan , Hong Yu

While Large Language Models (LLMs) have demonstrated high proficiency on English-centric medical examinations, their performance often declines when faced with non-English languages and multimodal diagnostic tasks. This study protocol…

One of the key goals of artificial intelligence (AI) is the development of a multimodal system that facilitates communication with the visual world (image and video) using a natural language query. Earlier works on medical question…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Deepak Gupta , Dina Demner-Fushman

LLMs hold great promise for healthcare applications, but the rapid evolution of medical knowledge and errors in training data often cause them to generate outdated or inaccurate information, limiting their applicability in high-stakes…

Computation and Language · Computer Science 2025-11-04 Shujun Xia , Haokun Lin , Yichen Wu , Yinan Zhou , Zixuan Li , Zhongwei Wan , Xingrun Xing , Yefeng Zheng , Xiang Li , Caifeng Shan , Zhenan Sun , Quanzheng Li

As the number of dementia patients rises, the need for accurate diagnostic procedures rises as well. Current methods, like using an MRI scan, rely on human input, which can be inaccurate. However, the decision logic behind machine learning…

Image and Video Processing · Electrical Eng. & Systems 2024-06-28 Tyler Morris , Ziming Liu , Longjian Liu , Xiaopeng Zhao

Medical artificial intelligence (AI) is revolutionizing the interpretation of chest X-ray (CXR) images by providing robust tools for disease diagnosis. However, the effectiveness of these AI models is often limited by their reliance on…

Image and Video Processing · Electrical Eng. & Systems 2024-10-14 Lijian Xu , Ziyu Ni , Hao Sun , Hongsheng Li , Shaoting Zhang

Medical diagnostic applications require models that can process multimodal medical inputs (images, patient histories, lab results) and generate diverse outputs including both textual reports and visual content (annotations, segmentation…

Surgery requires comprehensive medical knowledge, visual assessment skills, and procedural expertise. While recent surgical AI models have focused on solving task-specific problems, there is a need for general-purpose systems that can…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Samuel Schmidgall , Joseph Cho , Cyril Zakka , William Hiesinger
‹ Prev 1 3 4 5 6 7 10 Next ›