中文
相关论文

相关论文: MedVersa: A Generalist Foundation Model for Medica…

200 篇论文

Artificial intelligence applied to retinal images offers significant potential for recognizing signs and symptoms of retinal conditions and expediting the diagnosis of eye diseases and systemic disorders. However, developing generalized…

图像与视频处理 · 电气工程与系统科学 2024-08-19 Boa Jang , Youngbin Ahn , Eun Kyung Choe , Chang Ki Yoon , Hyuk Jin Choi , Young-Gon Kim

We present MedASR, an open-source 105M-parameter model engineered for high-accuracy medical dictation. Prioritizing a "small, fast, and accurate" design, MedASR addresses 3 core pillars (1) Data: overcoming clinical corpora scarcity and…

音频与语音处理 · 电气工程与系统科学 2026-05-19 Ke Wu , Ehsan Variani , Tom Bagby , Shashir Reddy , Rory Pilgrim

Deformable registration is a fundamental task in medical image processing, aiming to achieve precise alignment by establishing nonlinear correspondences between images. Traditional methods offer good adaptability and interpretability but…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Jing Hu , Kaiwei Yu , Hongjiang Xian , Shu Hu , Xin Wang

With the rapid growth of large language models (LLMs) and vision-language models (VLMs) in medicine, simply integrating clinical text and medical imaging does not guarantee reliable reasoning. Existing multimodal models often produce…

人工智能 · 计算机科学 2025-12-29 Zelin Zang , Wenyi Gu , Siqi Ma , Dan Yang , Yue Shen , Zhu Zhang , Guohui Fan , Wing-Kuen Ling , Fuji Yang

Building multisensory AI systems that learn from multiple sensory inputs such as text, speech, video, real-world sensors, wearable devices, and medical data holds great promise for impact in many scientific areas with practical benefits,…

机器学习 · 计算机科学 2024-05-01 Paul Pu Liang

The increasing use of tools and solutions based on Large Language Models (LLMs) for various tasks in the medical domain has become a prominent trend. Their use in this highly critical and sensitive domain has thus raised important questions…

Large language models (LLMs) have shown incredible proficiency in performing tasks that require semantic understanding of natural language instructions. Recently, many works have further expanded this capability to perceive multimodal audio…

Multi-modal large language models (MLLMs) have shown impressive capabilities as a general-purpose interface for various visual and linguistic tasks. However, building a unified MLLM for multi-task learning in the medical field remains a…

计算机视觉与模式识别 · 计算机科学 2024-11-04 Xun Zhu , Ying Hu , Fanbin Mo , Miao Li , Ji Wu

Despite remarkable advancements in pixel-level medical image perception, existing methods are either limited to specific tasks or heavily rely on accurate bounding boxes or text labels as input prompts. However, the medical knowledge…

计算机视觉与模式识别 · 计算机科学 2025-04-24 Qinyue Tong , Ziqian Lu , Jun Liu , Yangming Zheng , Zheming Lu

Artificial intelligence (AI)-enabled diagnostics in maxillofacial pathology require structured, high-quality multimodal datasets. However, existing resources provide limited ameloblastoma coverage and lack the format consistency needed for…

人工智能 · 计算机科学 2026-02-06 Ajo Babu George , Anna Mariam John , Athul Anoop , Balu Bhasuran

This paper proposes a MedGemma-based framework for automatic abnormality detection in musculoskeletal radiographs. Departing from conventional autoencoder and neural network pipelines, the proposed method leverages the MedGemma foundation…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Soumyajit Maity , Pranjal Kamboj , Sneha Maity , Rajat Singh , Sankhadeep Chatterjee

The integration of medical imaging and clinical text has enabled the emergence of generalist artificial intelligence (AI) systems for healthcare. However, pervasive biases, such as imbalanced disease prevalence, skewed anatomical region…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Cheng Li , Weijian Huang , Jiarun Liu , Hao Yang , Qi Yang , Song Wu , Ye Li , Hairong Zheng , Shanshan Wang

Frontier artificial intelligence (AI) models, such as OpenAI's GPT-5 and Meta's DINOv3, have advanced rapidly through training on internet-scale public data, yet such systems lack access to private clinical data. Neuroimaging, in…

The remarkable performance of the Transformer architecture in natural language processing has recently also triggered broad interest in Computer Vision. Among other merits, Transformers are witnessed as capable of learning long-range…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Reza Azad , Amirhossein Kazerouni , Moein Heidari , Ehsan Khodapanah Aghdam , Amirali Molaei , Yiwei Jia , Abin Jose , Rijo Roy , Dorit Merhof

Multimodal/vision language models (VLMs) are increasingly being deployed in healthcare settings worldwide, necessitating robust benchmarks to ensure their safety, efficacy, and fairness. Multiple-choice question and answer (QA) datasets…

Many clinical tasks require an understanding of specialized data, such as medical images and genomics, which is not typically found in general-purpose large multimodal models. Building upon Gemini's multimodal models, we develop several…

A multitude of work has shown that machine learning-based medical diagnosis systems can be biased against certain subgroups of people. This has motivated a growing number of bias mitigation algorithms that aim to address fairness issues in…

机器学习 · 计算机科学 2023-02-21 Yongshuo Zong , Yongxin Yang , Timothy Hospedales

Dermatological care via telemedicine often lacks the rich context of in-person visits. Clinicians must make diagnoses based on a handful of images and brief descriptions, without the benefit of physical exams, second opinions, or reference…

人工智能 · 计算机科学 2025-08-27 Karishma Thakrar , Shreyas Basavatia , Akshay Daftardar

Despite tremendous progress in computer vision, there has not been an attempt for machine learning on very large-scale medical image databases. We present an interleaved text/image deep learning system to extract and mine the semantic…

计算机视觉与模式识别 · 计算机科学 2015-05-05 Hoo-Chang Shin , Le Lu , Lauren Kim , Ari Seff , Jianhua Yao , Ronald M. Summers