中文
相关论文

相关论文: A Foundational Multimodal Vision Language AI Assis…

200 篇论文

The capabilities of AI for biomedicine span a wide spectrum, from the atomic level, where it solves partial differential equations for quantum systems, to the molecular level, predicting chemical or protein structures, and further extending…

计算与语言 · 计算机科学 2024-03-26 Zhenyu Bi , Sajib Acharjee Dip , Daniel Hajialigol , Sindhura Kommu , Hanwen Liu , Meng Lu , Xuan Wang

In healthcare and medical diagnostics, Visual Question Answering (VQA) mayemergeasapivotal tool in scenarios where analysis of intricate medical images becomes critical for accurate diagnoses. Current text-based VQA systems limit their…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Tonmoy Rajkhowa , Amartya Roy Chowdhury , Sankalp Nagaonkar , Achyut Mani Tripathi

Developing generalist foundation model has recently attracted tremendous attention among researchers in the field of AI for Medicine (AI4Medicine). A pivotal insight in developing these models is their reliance on dataset scaling, which…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Xiaoman Zhang , Chaoyi Wu , Ziheng Zhao , Jiayu Lei , Ya Zhang , Yanfeng Wang , Weidi Xie

Addressing the gap in understanding visual comprehension in Large Language Models (LLMs), we designed a challenge-response study, subjecting Google Bard and GPT-Vision to 64 visual tasks, spanning categories like "Visual Situational…

计算机视觉与模式识别 · 计算机科学 2023-10-18 David Noever , Samantha Elizabeth Miller Noever

The transition from task-specific artificial intelligence toward general-purpose foundation models raises fundamental questions about their capacity to support the integrated reasoning required in clinical medicine, where diagnosis demands…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Alexandru Florea , Shansong Wang , Mingzhe Hu , Qiang Li , Zach Eidex , Luke del Balzo , Mojtaba Safari , Xiaofeng Yang

Documentation burden is a major contributor to clinician burnout, which is rising nationally and is an urgent threat to our ability to care for patients. Artificial intelligence (AI) chatbots, such as ChatGPT, could reduce clinician burden…

Artificial intelligence (AI)-enabled diagnostics in maxillofacial pathology require structured, high-quality multimodal datasets. However, existing resources provide limited ameloblastoma coverage and lack the format consistency needed for…

人工智能 · 计算机科学 2026-02-06 Ajo Babu George , Anna Mariam John , Athul Anoop , Balu Bhasuran

Inpatient pathways demand complex clinical decision-making based on comprehensive patient information, posing critical challenges for clinicians. Despite advancements in large language models (LLMs) in medical applications, limited research…

人工智能 · 计算机科学 2025-03-18 Zhen Chen , Zhihao Peng , Xusheng Liang , Cheng Wang , Peigan Liang , Linsheng Zeng , Minjie Ju , Yixuan Yuan

Recently large vision-language models have shown potential when interpreting complex images and generating natural language descriptions using advanced reasoning. Medicine's inherently multimodal nature incorporating scans and text-based…

图像与视频处理 · 电气工程与系统科学 2024-07-15 Naman Sharma

Pathological structures in medical images are typically deviations from the expected anatomy of a patient. While clinicians consider this interplay between anatomy and pathology, recent deep learning algorithms specialize in recognizing…

The rapid development of Multimodal Large Language Models (MLLMs), such as GPT-4o, marks a significant step toward artificial general intelligence. Existing methods typically align vision encoders with LLMs via supervised fine-tuning (SFT),…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Hai-Long Sun , Da-Wei Zhou , Yang Li , Shiyin Lu , Chao Yi , Qing-Guo Chen , Zhao Xu , Weihua Luo , Kaifu Zhang , De-Chuan Zhan , Han-Jia Ye

Background/Objectives: Age-related macular degeneration, glaucoma, diabetic retinopathy (DR), diabetic macular edema, and pathological myopia affect hundreds of millions of people worldwide. Early screening for these diseases is essential,…

图像与视频处理 · 电气工程与系统科学 2025-07-18 Ananya Raghu , Anisha Raghu , Alice S. Tang , Yannis M. Paulus , Tyson N. Kim , Tomiko T. Oskotsky

Pathology image segmentation across multiple centers encounters significant challenges due to diverse sources of heterogeneity including imaging modalities, organs, and scanning equipment, whose variability brings representation bias and…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Yuan Zhang , Feng Chen , Yaolei Qi , Guanyu Yang , Huazhu Fu

Several pathologies can alter the way people walk, i.e. their gait. Gait analysis can therefore be used to detect impairments and help diagnose illnesses and assess patient recovery. Using vision-based systems, diagnoses could be done at…

计算机视觉与模式识别 · 计算机科学 2021-10-07 Pedro Albuquerque , Joao Machado , Tanmay Tulsidas Verlekar , Luis Ducla Soares , Paulo Lobato Correia

Artificial intelligence (AI) methodologies hold great promise for the rapid and accurate diagnosis of coronary artery disease (CAD) from intravascular optical coherent tomography (IVOCT) images. Numerous papers have been published…

图像与视频处理 · 电气工程与系统科学 2025-02-03 Xu Chen , Yuan Huang , Benn Jessney , Jason Sangha , Sophie Gu , Carola-Bibiane Schönlieb , Martin Bennett , Michael Roberts

This work conducts an evaluation of GPT-4V's multimodal capability for medical image analysis, with a focus on three representative tasks of radiology report generation, medical visual question answering, and medical visual grounding. For…

计算机视觉与模式识别 · 计算机科学 2024-02-01 Yingshu Li , Yunyi Liu , Zhanyu Wang , Xinyu Liang , Lei Wang , Lingqiao Liu , Leyang Cui , Zhaopeng Tu , Longyue Wang , Luping Zhou

ChatGPT is an advanced natural language processing tool with growing applications across various disciplines in medical research. Thematic analysis, a qualitative research method to identify and interpret patterns in data, is one…

计算与语言 · 计算机科学 2025-05-08 V Vien Lee , Stephanie C. C. van der Lubbe , Lay Hoon Goh , Jose M. Valderas

Automatically evaluating vision-language tasks is challenging, especially when it comes to reflecting human judgments due to limitations in accounting for fine-grained details. Although GPT-4V has shown promising results in various…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Xinlu Zhang , Yujie Lu , Weizhi Wang , An Yan , Jun Yan , Lianke Qin , Heng Wang , Xifeng Yan , William Yang Wang , Linda Ruth Petzold

Automating medical report generation from histopathology images is a critical challenge requiring effective visual representations and domain-specific knowledge. Inspired by the common practices of human experts, we propose an in-context…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Shih-Wen Liu , Hsuan-Yu Fan , Wei-Ta Chu , Fu-En Yang , Yu-Chiang Frank Wang

Large multimodal language models (LMMs) have achieved significant success in general domains. However, due to the significant differences between medical images and text and general web content, the performance of LMMs in medical scenarios…

计算机视觉与模式识别 · 计算机科学 2023-06-23 Weihao Gao , Zhuo Deng , Zhiyuan Niu , Fuju Rong , Chucheng Chen , Zheng Gong , Wenze Zhang , Daimin Xiao , Fang Li , Zhenjie Cao , Zhaoyi Ma , Wenbin Wei , Lan Ma