中文
相关论文

相关论文: Navigating Gigapixel Pathology Images with Large M…

200 篇论文

Multimodal foundation models have shown compelling but conflicting performance in medical image interpretation. However, the mechanisms by which these models integrate and prioritize different data modalities, including images and text,…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Thomas Buckley , James A. Diao , Pranav Rajpurkar , Adam Rodman , Arjun K. Manrai

Vision Language Models (VLMs) like CLIP have attracted substantial attention in pathology, serving as backbones for applications such as zero-shot image classification and Whole Slide Image (WSI) analysis. Additionally, they can function as…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Yuxuan Sun , Yunlong Zhang , Yixuan Si , Chenglu Zhu , Zhongyi Shui , Kai Zhang , Jingxiong Li , Xingheng Lyu , Tao Lin , Lin Yang

High-quality textual training data is essential for the success of multimodal data processing tasks, yet outputs from image captioning models like BLIP and GIT often contain errors and anomalies that are difficult to rectify using…

计算与语言 · 计算机科学 2025-02-25 Elyas Meguellati , Nardiena Pratama , Shazia Sadiq , Gianluca Demartini

Learning good representation of giga-pixel level whole slide pathology images (WSI) for downstream tasks is critical. Previous studies employ multiple instance learning (MIL) to represent WSIs as bags of sampled patches because, for most…

计算机视觉与模式识别 · 计算机科学 2022-12-01 Chunyuan Li , Xinliang Zhu , Jiawen Yao , Junzhou Huang

ChatGPT explores a strategic blueprint of question answering (QA) in delivering medical diagnosis, treatment recommendations, and other healthcare support. This is achieved through the increasing incorporation of medical domain data via…

计算与语言 · 计算机科学 2024-01-23 Qing Li , Lei Li , Yu Li

Large language models (LLMs) can simulate clinical reasoning based on natural language prompts, but their utility in ophthalmology is largely unexplored. This study evaluated GPT-4's ability to interpret structured textual descriptions of…

Due to the large size and lack of fine-grained annotation, Whole Slide Images (WSIs) analysis is commonly approached as a Multiple Instance Learning (MIL) problem. However, previous studies only learn from training data, posing a stark…

计算机视觉与模式识别 · 计算机科学 2024-11-28 Weiqin Zhao , Ziyu Guo , Yinshuang Fan , Yuming Jiang , Maximus Yeung , Lequan Yu

Passively collected behavioral health data from ubiquitous sensors holds significant promise to provide mental health professionals insights from patient's daily lives; however, developing analysis tools to use this data in clinical…

Large Multimodal Models (LMMs) have achieved impressive success in visual understanding and reasoning, remarkably improving the performance of mathematical reasoning in a visual context. Yet, a challenging type of visual math lies in the…

计算机视觉与模式识别 · 计算机科学 2024-05-09 Yunxin Li , Baotian Hu , Haoyuan Shi , Wei Wang , Longyue Wang , Min Zhang

Geolocation, the task of identifying the geographic location of an image, requires abundant world knowledge and complex reasoning abilities. Though advanced large multimodal models (LMMs) have shown superior aforementioned capabilities,…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Yushuo Zheng , Huiyu Duan , Zicheng Zhang , Xiaohong Liu , Xiongkuo Min

Backgr: Digital pathology images are increasingly used both for diagnosis and research, because slide scanners are nowadays broadly available and because the quantitative study of these images yields new insights in systems biology.…

定量方法 · 定量生物学 2017-09-08 Christophe Deroulers , David Ameisen , Mathilde Badoual , Chloé Gerin , Alexandre Granier , Marc Lartaud

Understanding the deep semantics of images is essential in the era dominated by social media. However, current research works primarily on the superficial description of images, revealing a notable deficiency in the systematic investigation…

计算与语言 · 计算机科学 2024-06-21 Yixin Yang , Zheng Li , Qingxiu Dong , Heming Xia , Zhifang Sui

While large multi-modal models (LMM) have shown notable progress in multi-modal tasks, their capabilities in tasks involving dense textual content remains to be fully explored. Dense text, which carries important information, is often found…

计算与语言 · 计算机科学 2024-05-14 Shuo Zhang , Biao Yang , Zhang Li , Zhiyin Ma , Yuliang Liu , Xiang Bai

In computational pathology, extracting spatial features from gigapixel whole slide images (WSIs) is a fundamental task, but due to their large size, WSIs are typically segmented into smaller tiles. A critical aspect of this analysis is…

图像与视频处理 · 电气工程与系统科学 2024-10-23 Ruiwen Ding , Kha-Dinh Luong , Erika Rodriguez , Ana Cristina Araujo Lemos da Silva , William Hsu

Prompt learning has demonstrated impressive efficacy in the fine-tuning of multimodal large models to a wide range of downstream tasks. Nonetheless, applying existing prompt learning methods for the diagnosis of neurological disorder still…

计算机视觉与模式识别 · 计算机科学 2024-06-28 Liang Peng , Songyue Cai , Zongqian Wu , Huifang Shang , Xiaofeng Zhu , Xiaoxiao Li

Documents are fundamental to preserving and disseminating information, often incorporating complex layouts, tables, and charts that pose significant challenges for automatic document understanding (DU). While vision-language large models…

计算与语言 · 计算机科学 2025-06-19 Negar Foroutan , Angelika Romanou , Matin Ansaripour , Julian Martin Eisenschlos , Karl Aberer , Rémi Lebret

Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understanding. However, existing benchmarks typically evaluate agents with fully synthetic, single-turn…

Scientific data visualization plays a crucial role in research by enabling the direct display of complex information and assisting researchers in identifying implicit patterns. Despite its importance, the use of Large Language Models (LLMs)…

Accurate diagnosis of skin diseases remains a significant challenge due to the complex and diverse visual features present in dermatoscopic images, often compounded by a lack of interpretability in existing purely visual diagnostic models.…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Kexin Yu , Zihan Xu , Jialei Xie , Carter Adams

Recent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image- and video-language understanding. Conversely, less attention…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Ye Liu , Zongyang Ma , Junfu Pu , Zhongang Qi , Yang Wu , Ying Shan , Chang Wen Chen