中文
相关论文

相关论文: A Foundational Multimodal Vision Language AI Assis…

200 篇论文

Despite the promises of data-driven artificial intelligence (AI), little is known about how we can bridge the gulf between traditional physician-driven diagnosis and a plausible future of medicine automated by AI. Specifically, how can we…

人机交互 · 计算机科学 2021-02-12 Hongyan Gu , Jingbin Huang , Lauren Hung , Xiang 'Anthony' Chen

Is it possible to develop an "AI Pathologist" to pass the board-certified examination of the American Board of Pathology? To achieve this goal, the first step is to create a visual question answering (VQA) dataset where the AI agent is…

计算与语言 · 计算机科学 2020-03-24 Xuehai He , Yichen Zhang , Luntian Mou , Eric Xing , Pengtao Xie

Despite the progress made by multimodal large language models (MLLMs) in computational pathology, they remain limited by a predominant focus on patch-level analysis, missing essential contextual information at the whole-slide level. The…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Ying Chen , Guoan Wang , Yuanfeng Ji , Yanjun Li , Jin Ye , Tianbin Li , Ming Hu , Rongshan Yu , Yu Qiao , Junjun He

Interpretability is significant in computational pathology, leading to the development of multimodal information integration from histopathological image and corresponding text data.However, existing multimodal methods have limited…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Kangcheng Zhou , Jun Jiang , Qing Zhang , Shuang Zheng , Qingli Li , Shugong Xu

Pathologists diagnose cancer using gigapixel whole-slide images (WSIs), but the current digital workflow is fragmented. These multiscale datasets often exceed 100,000 x 100,000 pixels, yet standard 2D monitors restrict the field of view.…

Whole-slide image (WSI) preprocessing, comprising tissue detection followed by patch extraction, is foundational to AI-driven computational pathology but remains a major bottleneck for scaling to large and heterogeneous cohorts. We present…

Recent advances in computational pathology have led to the emergence of numerous foundation models. These models typically rely on general-purpose encoders with multi-instance learning for whole slide image (WSI) classification or apply…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Yuxuan Sun , Yixuan Si , Chenglu Zhu , Kai Zhang , Zhongyi Shui , Bowen Ding , Tao Lin , Lin Yang

Driven by the large foundation models, the development of artificial intelligence has witnessed tremendous progress lately, leading to a surge of general interest from the public. In this study, we aim to assess the performance of OpenAI's…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Chaoyi Wu , Jiayu Lei , Qiaoyu Zheng , Weike Zhao , Weixiong Lin , Xiaoman Zhang , Xiao Zhou , Ziheng Zhao , Ya Zhang , Yanfeng Wang , Weidi Xie

Recent advances in vision language models (VLMs) have enabled broad progress in the general medical field. However, pathology still remains a more challenging subdomain, with current pathology specific VLMs exhibiting limitations in both…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Wenchuan Zhang , Penghao Zhang , Jingru Guo , Tao Cheng , Jie Chen , Shuwan Zhang , Zhang Zhang , Yuhao Yi , Hong Bu

Is it possible to develop an "AI Pathologist" to pass the board-certified examination of the American Board of Pathology (ABP)? To build such a system, three challenges need to be addressed. First, we need to create a visual question…

计算机视觉与模式识别 · 计算机科学 2020-10-26 Xuehai He , Zhuo Cai , Wenlan Wei , Yichen Zhang , Luntian Mou , Eric Xing , Pengtao Xie

We present VisionFM, a foundation model pre-trained with 3.4 million ophthalmic images from 560,457 individuals, covering a broad range of ophthalmic diseases, modalities, imaging devices, and demography. After pre-training, VisionFM…

Rare cancers comprise 20-25% of all malignancies but face major diagnostic challenges due to limited expert availability-especially in pediatric oncology, where they represent over 70% of cases. While pathology vision-language (VL)…

计算机视觉与模式识别 · 计算机科学 2025-08-25 Dexuan He , Xiao Zhou , Wenbin Guan , Liyuan Zhang , Xiaoman Zhang , Sinuo Xu , Ge Wang , Lifeng Wang , Xiaojun Yuan , Xin Sun , Yanfeng Wang , Kun Sun , Ya Zhang , Weidi Xie

With the development of generative artificial intelligence and instruction tuning techniques, multimodal large language models (MLLMs) have made impressive progress on general reasoning tasks. Benefiting from the chain-of-thought (CoT)…

机器学习 · 计算机科学 2025-07-03 Junjie Zhou , Yingli Zuo , Shichang Feng , Peng Wan , Qi Zhu , Daoqiang Zhang , Wei Shao

In this paper, we present a large-scale evaluation probing GPT-4V's capabilities and limitations for biomedical image analysis. GPT-4V represents a breakthrough in artificial general intelligence (AGI) for computer vision, with applications…

Pathological diagnosis remains the definitive standard for identifying tumors. The rise of multimodal large models has simplified the process of integrating image analysis with textual descriptions. Despite this advancement, the substantial…

计算机视觉与模式识别 · 计算机科学 2024-08-14 Xiaomin Wu , Rui Xu , Pengchen Wei , Wenkang Qin , Peixiang Huang , Ziheng Li , Lin Luo

Pathology image segmentation is crucial in computational pathology for analyzing histological features relevant to cancer diagnosis and prognosis. However, current methods face major challenges in clinical applications due to limited…

计算机视觉与模式识别 · 计算机科学 2025-08-20 Zhixuan Chen , Junlin Hou , Liqi Lin , Yihui Wang , Yequan Bie , Xi Wang , Yanning Zhou , Ronald Cheong Kin Chan , Hao Chen

The application of artificial intelligence (AI) in IVF has shown promise in improving consistency and standardization of decisions, but often relies on annotated data and does not make use of the multimodal nature of IVF data. We…

[18F]FDG-PET/CT is a cornerstone imaging modality for tumor staging and treatment response assessment across many cancer types, yet expert reader shortages necessitate more efficient diagnostic aids. While standalone AI models for automatic…

Medical AI assistants support doctors in disease diagnosis, medical image analysis, and report generation. However, they still face significant challenges in clinical use, including limited accuracy with multimodal content and insufficient…

计算机视觉与模式识别 · 计算机科学 2025-05-07 Haonan Wang , Jiaji Mao , Lehan Wang , Qixiang Zhang , Marawan Elbatel , Yi Qin , Huijun Hu , Baoxun Li , Wenhui Deng , Weifeng Qin , Hongrui Li , Jialin Liang , Jun Shen , Xiaomeng Li

Accurate analysis of pathological images is essential for automated tumor diagnosis but remains challenging due to high structural similarity and subtle morphological variations in tissue images. Current vision-language (VL) models often…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Yating Huang , Ziyan Huang , Lintao Xiang , Qijun Yang , Hujun Yin