中文
相关论文

相关论文: A Foundational Multimodal Vision Language AI Assis…

200 篇论文

The advancement of artificial intelligence (AI) for organ segmentation and tumor detection is propelled by the growing availability of computed tomography (CT) datasets with detailed, per-voxel annotations. However, these AI models often…

图像与视频处理 · 电气工程与系统科学 2024-05-29 Jie Liu , Yixiao Zhang , Kang Wang , Mehmet Can Yavuz , Xiaoxi Chen , Yixuan Yuan , Haoliang Li , Yang Yang , Alan Yuille , Yucheng Tang , Zongwei Zhou

Multimodal large models have shown great potential in automating pathology image analysis. However, current multimodal models for gastrointestinal pathology are constrained by both data quality and reasoning transparency: pervasive noise…

图像与视频处理 · 电气工程与系统科学 2025-07-25 Minxi Ouyang , Lianghui Zhu , Yaqing Bao , Qiang Huang , Jingli Ouyang , Tian Guan , Xitong Ling , Jiawen Li , Song Duan , Wenbin Dai , Li Zheng , Xuemei Zhang , Yonghong He

Foundation models and vision-language pre-training have significantly advanced Vision-Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their application in domain-specific agricultural tasks,…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Khang Nguyen Quoc , Phuong D. Dao , Luyl-Da Quach

The integration of deep learning-based glaucoma detection with large language models (LLMs) presents an automated strategy to mitigate ophthalmologist shortages and improve clinical reporting efficiency. However, applying general LLMs to…

多智能体系统 · 计算机科学 2025-12-18 Philip R. Liu , Sparsh Bansal , Jimmy Dinh , Aditya Pawar , Ramani Satishkumar , Shail Desai , Neeraj Gupta , Xin Wang , Shu Hu

Multimodal artificial intelligence (AI) systems have the potential to enhance clinical decision-making by interpreting various types of medical data. However, the effectiveness of these models across all medical fields is uncertain. Each…

Artificial intelligence (AI) has significant potential in healthcare applications, but its training and deployment faces challenges due to healthcare's diverse data, complex tasks, and the need to preserve privacy. Foundation models that…

Forensic pathology is critical in analyzing death manner and time from the microscopic aspect to assist in the establishment of reliable factual bases for criminal investigation. In practice, even the manual differentiation between…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Chen Shen , Jun Zhang , Xinggong Liang , Zeyi Hao , Kehan Li , Fan Wang , Zhenyuan Wang , Chunfeng Lian

The diagnosis of pathological images is often limited by expert availability and regional disparities, highlighting the importance of automated diagnosis using Vision-Language Models (VLMs). Traditional multimodal models typically emphasize…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Jianyu Wu , Hao Yang , Xinhua Zeng , Guibing He , Zhiyu Chen , Zihui Li , Xiaochuan Zhang , Yangyang Ma , Run Fang , Yang Liu

Artificial Intelligence (AI) has revolutionized various fields, including medicine and mental health support. One promising application is ChatGPT, an advanced conversational AI model that uses deep learning techniques to provide human-like…

神经元与认知 · 定量生物学 2023-11-16 Farzan Vahedifard , Atieh Sadeghniiat Haghighi , Tirth Dave , Mohammad Tolouei , Fateme Hoshyar Zare

Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains,…

Recent advances in AI combine large language models (LLMs) with vision encoders that bring forward unprecedented technical capabilities to leverage for a wide range of healthcare applications. Focusing on the domain of radiology,…

Recently, Multimodal Large Language Models (MLLMs) have gained significant attention for their remarkable ability to process and analyze non-textual data, such as images, videos, and audio. Notably, several adaptations of general-domain…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Wenhui Zhu , Xin Li , Xiwen Chen , Peijie Qiu , Vamsi Krishna Vasa , Xuanzhao Dong , Yanxi Chen , Natasha Lepore , Oana Dumitrascu , Yi Su , Yalin Wang

Widespread clinical deployment of computer-aided diagnosis (CAD) systems is hindered by the challenge of integrating with existing hospital IT infrastructure. Here, we introduce VisionCAD, a vision-based radiological assistance framework…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Jiaming Li , Junlei Wu , Sheng Wang , Honglin Xiong , Jiangdong Cai , Zihao Zhao , Yitao Zhu , Yuan Yin , Dinggang Shen , Qian Wang

Virtual immunohistochemistry (IHC) aims to computationally synthesize molecular staining patterns from routine Hematoxylin and Eosin (H\&E) images, offering a cost-effective and tissue-efficient alternative to traditional physical staining.…

计算机视觉与模式识别 · 计算机科学 2026-01-23 Rongze Ma , Mengkang Lu , Zhenyu Xiang , Yongsheng Pan , Yicheng Wu , Qingjie Zeng , Yong Xia

To develop intelligent speech assistants and integrate them seamlessly with intra-operative decision-support frameworks, accurate and efficient surgical phase recognition is a prerequisite. In this study, we propose a multimodal framework…

音频与语音处理 · 电气工程与系统科学 2024-06-24 Kubilay Can Demir , Belen Lojo Rodriguez , Tobias Weise , Andreas Maier , Seung Hee Yang

Automatic pathological speech detection approaches have shown promising results, gaining attention as potential diagnostic tools alongside costly traditional methods. While these approaches can achieve high accuracy, their lack of…

音频与语音处理 · 电气工程与系统科学 2025-04-01 Mahdi Amiri , Hatef Otroshi Shahreza , Ina Kodrasi

Generative artificial intelligence (AI) models, such as diffusion models and OpenAI's ChatGPT, are transforming medicine by enhancing diagnostic accuracy and automating clinical workflows. The field has advanced rapidly, evolving from…

人工智能 · 计算机科学 2025-08-28 Lukas Buess , Matthias Keicher , Nassir Navab , Andreas Maier , Soroosh Tayebi Arasteh

AI-assisted imaging made substantial advances in tumor diagnosis and management. However, a major barrier to developing robust oncology foundation models is the scarcity of large-scale, high-quality annotated datasets, which are limited by…

The integration of medical imaging and clinical text has enabled the emergence of generalist artificial intelligence (AI) systems for healthcare. However, pervasive biases, such as imbalanced disease prevalence, skewed anatomical region…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Cheng Li , Weijian Huang , Jiarun Liu , Hao Yang , Qi Yang , Song Wu , Ye Li , Hairong Zheng , Shanshan Wang

Computational pathology demands both visual pattern recognition and dynamic integration of structured domain knowledge, including taxonomy, grading criteria, and clinical evidence. In practice, diagnostic reasoning requires linking…