中文
相关论文

相关论文: A Foundational Multimodal Vision Language AI Assis…

200 篇论文

Quantitative evaluation of echocardiography is essential for precise assessment of cardiac condition, monitoring disease progression, and guiding treatment decisions. The diverse nature of echo images, including variations in probe types,…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Abdoul Aziz Amadou , Yue Zhang , Sebastien Piat , Paul Klein , Ingo Schmuecking , Tiziano Passerini , Puneet Sharma

Recent rapid progress in the field of computational pathology has been enabled by foundation models. These models are beginning to move beyond encoding image patches towards whole-slide understanding but their clinical utility remains…

Despite the promise of computational pathology foundation models, adapting them to specific clinical tasks remains challenging due to the complexity of whole-slide image (WSI) processing, the opacity of learned features, and the wide range…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Abdul Rahman Diab , Emily E. Karn , Renchin Wu , Emily S. Ruiz , William Lotter

Deep learning has been increasingly incorporated into various computational pathology applications to improve its efficiency, accuracy, and robustness. Although successful, most previous approaches for image classification have crucial…

图像与视频处理 · 电气工程与系统科学 2024-07-15 Anh Tien Nguyen , Jin Tae Kwak

Biomedical multimodal assistants have the potential to unify radiology, pathology, and clinical-text reasoning, yet a critical deployment gap remains: top-performing systems are either closed-source or computationally prohibitive,…

In the current landscape of artificial intelligence, foundation models serve as the bedrock for advancements in both language and vision domains. OpenAI GPT-4 has emerged as the pinnacle in large language models (LLMs), while the computer…

计算机视觉与模式识别 · 计算机科学 2023-11-20 Chris Kelly , Luhui Hu , Cindy Yang , Yu Tian , Deshun Yang , Bang Yang , Zaoshan Huang , Zihao Li , Yuexian Zou

Medical image classification requires labeled, task-specific datasets which are used to train deep learning networks de novo, or to fine-tune foundation models. However, this process is computationally and technically demanding. In language…

The evolution of text to visual components facilitates people's daily lives, such as generating image, videos from text and identifying the desired elements within the images. Computer vision models involving the multimodal abilities in the…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Chris Kelly , Luhui Hu , Jiayin Hu , Yu Tian , Deshun Yang , Bang Yang , Cindy Yang , Zihao Li , Zaoshan Huang , Yuexian Zou

Faces and humans are crucial elements in social interaction and are widely included in everyday photos and videos. Therefore, a deep understanding of faces and humans will enable multi-modal assistants to achieve improved response quality…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Lixiong Qin , Shilong Ou , Miaoxuan Zhang , Jiangning Wei , Yuhang Zhang , Xiaoshuai Song , Yuchen Liu , Mei Wang , Weiran Xu

We present Singpath-VL, a vision-language large model, to fill the vacancy of AI assistant in cervical cytology. Recent advances in multi-modal large language models (MLLMs) have significantly propelled the field of computational pathology.…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Zhen Qiu , Kaiwen Xiao , Zhengwei Lu , Xiangyu Liu , Lei Zhao , Hao Zhang

Gastrointestinal (GI) diseases represent a clinically significant burden, necessitating precise diagnostic approaches to optimize patient outcomes. Conventional histopathological diagnosis suffers from limited reproducibility and diagnostic…

While Vision-Language Models (VLMs) have achieved notable progress in computational pathology (CPath), the gigapixel scale and spatial heterogeneity of Whole Slide Images (WSIs) continue to pose challenges for multimodal understanding.…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Fengchun Liu , Songhan Jiang , Linghan Cai , Ziyue Wang , Yongbing Zhang

Chest X-ray interpretation is one of the most frequently performed diagnostic tasks in medicine and a primary target for AI development, yet current vision-language models are primarily trained on datasets of paired images and reports, not…

The semantic segmentation task in pathology plays an indispensable role in assisting physicians in determining the condition of tissue lesions. With the proposal of Segment Anything Model (SAM), more and more foundation models have seen…

图像与视频处理 · 电气工程与系统科学 2024-09-05 Mingya Zhang , Liang Wang , Zhihao Chen , Yiyuan Ge , Xianping Tao

The prediction of chemical synthesis pathways plays a pivotal role in materials science research. Challenges, such as the complexity of synthesis pathways and the lack of comprehensive datasets, currently hinder our ability to predict these…

材料科学 · 物理学 2023-11-03 Ziyi Chen , Fankai Xie , Meng Wan , Yang Yuan , Miao Liu , Zongguo Wang , Sheng Meng , Yangang Wang

Advances in artificial intelligence (AI) have achieved expert-level performance in medical imaging applications. Notably, self-supervised vision-language foundation models can detect a broad spectrum of pathologies without relying on…

计算机与社会 · 计算机科学 2024-02-23 Yuzhe Yang , Yujia Liu , Xin Liu , Avanti Gulhane , Domenico Mastrodicasa , Wei Wu , Edward J Wang , Dushyant W Sahani , Shwetak Patel

OpenAI released version GPT-4 on March 14, 2023, following the success of ChatGPT, which was announced in November 2022. In addition to the existing GPT-3 features, GPT-4 can interpret images. To achieve this, the processing power and model…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Omer Aydin , Enis Karaarslan

Vision Language Models (VLMs) like CLIP have attracted substantial attention in pathology, serving as backbones for applications such as zero-shot image classification and Whole Slide Image (WSI) analysis. Additionally, they can function as…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Yuxuan Sun , Yunlong Zhang , Yixuan Si , Chenglu Zhu , Zhongyi Shui , Kai Zhang , Jingxiong Li , Xingheng Lyu , Tao Lin , Lin Yang

Artificial intelligence (AI) and machine learning have changed the nature of scientific inquiry in recent years. Of these, the development of virtual assistants has accelerated greatly in the past few years, with ChatGPT becoming a…

计算机与社会 · 计算机科学 2023-05-25 Sukhpal Singh Gill , Rupinder Kaur

Visual search is important in our daily life. The efficient allocation of visual attention is critical to effectively complete visual search tasks. Prior research has predominantly modelled the spatial allocation of visual attention in…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Yini Fang , Jingling Yu , Haozheng Zhang , Ralf van der Lans , Bertram Shi