English
Related papers

Related papers: Navigating Gigapixel Pathology Images with Large M…

200 papers

Multimodal Large Language Models (MLLMs) are increasingly applied in real-world scenarios where user-provided images are often imperfect, requiring active image manipulations such as cropping, editing, or enhancement to uncover salient…

Accurate classification of pediatric central nervous system tumors remains challenging due to histological complexity and limited training data. While pathology foundation models have advanced whole-slide image (WSI) analysis, they often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Jian Yu , Joakim Nguyen , Jinrui Fang , Awais Naeem , Zeyuan Cao , Sanjay Krishnan , Nicholas Konz , Tianlong Chen , Chandra Krishnan , Hairong Wang , Edward Castillo , Ying Ding , Ankita Shukla

The scaling laws and extraordinary performance of large foundation models motivate the development and utilization of such models in biomedicine. However, despite early promising results on some biomedical benchmarks, there are still major…

Accurate analysis of histopathological images is critical for disease diagnosis and treatment planning. Whole-slide images (WSIs), which digitize tissue specimens at gigapixel resolution, are fundamental to this process but require…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Enhui Chai , Sicheng Chen , Tianyi Zhang , Chad Wong , Kecheng Huang , Zeyu Liu , Fei Xia

Whole Slide Images (WSIs) exhibit hierarchical structure, where diagnostic information emerges from cellular morphology, regional tissue organization, and global context. Existing Computational Pathology (CPath) Multimodal Large Language…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Basit Alawode , Arif Mahmood , Muaz Khalifa Al-Radi , Shahad Albastaki , Asim Khan , Muhammad Bilal , Moshira Ali Abdalla , Mohammed Bennamoun , Sajid Javed

Whole-slide images (WSIs) are an important data modality in computational pathology, yet their gigapixel resolution and lack of fine-grained annotations challenge conventional deep learning models. Multiple instance learning (MIL) offers a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Qian Zeng , Yihui Wang , Shu Yang , Yingxue Xu , Fengtao Zhou , Jiabo Ma , Dejia Cai , Zhengyu Zhang , Lijuan Qu , Yu Wang , Li Liang , Hao Chen

Medical Visual Question Answering (VQA) enhances clinical decision-making by enabling systems to interpret medical images and answer clinical queries. However, developing efficient, high-performance VQA models is challenging due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Belal Alsinglawi , Chris McCarthy , Sara Webb , Christopher Fluke , Navid Toosy Saidy

Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of…

Recent advancements in artificial intelligence (AI) have precipitated significant breakthroughs in healthcare, particularly in refining diagnostic procedures. However, previous studies have often been constrained to limited functionalities.…

Artificial Intelligence · Computer Science 2024-07-08 Asma Alkhaldi , Raneem Alnajim , Layan Alabdullatef , Rawan Alyahya , Jun Chen , Deyao Zhu , Ahmed Alsinan , Mohamed Elhoseiny

The rapidly emerging field of deep learning-based computational pathology has shown promising results in utilizing whole slide images (WSIs) to objectively prognosticate cancer patients. However, most prognostic methods are currently…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Mingxin Liu , Yunzan Liu , Hui Cui , Chunquan Li , Jiquan Ma

Visual spatial intelligence is critical for medical image interpretation, yet remains largely unexplored in Multimodal Large Language Models (MLLMs) for 3D imaging. This gap persists due to a systemic lack of datasets featuring structured…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Quoc-Huy Trinh , Xi Ding , Yang Liu , Zhenyue Qin , Xingjian Li , Gorkem Durak , Halil Ertugrul Aktas , Elif Keles , Ulas Bagci , Min Xu

We introduce VisualQuest, a novel dataset designed to rigorously evaluate multimodal large language models (MLLMs) on abstract visual reasoning tasks that require the integration of symbolic, cultural, and linguistic knowledge. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Kelaiti Xiao , Liang Yang , Dongyu Zhang , Paerhati Tulajiang , Hongfei Lin

AI tools in pathology have improved screening throughput, standardized quantification, and revealed prognostic patterns that inform treatment. However, adoption remains limited because most systems still lack the human-readable reasoning…

Artificial Intelligence · Computer Science 2025-11-18 Yunqi Hong , Johnson Kao , Liam Edwards , Nein-Tzu Liu , Chung-Yen Huang , Alex Oliveira-Kowaleski , Cho-Jui Hsieh , Neil Y. C. Lin

Pathology reports are rich in clinical and pathological details but are often presented in free-text format. The unstructured nature of these reports presents a significant challenge limiting the accessibility of their content. In this…

Computation and Language · Computer Science 2024-11-28 Ethar Alzaid , Gabriele Pergola , Harriet Evans , David Snead , Fayyaz Minhas

Medical large vision-language models (LVLMs) have demonstrated promising performance across various single-image question answering (QA) benchmarks, yet their capability in processing multi-image clinical scenarios remains underexplored.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Xikai Yang , Juzheng Miao , Yuchen Yuan , Jiaze Wang , Qi Dou , Jinpeng Li , Pheng-Ann Heng

Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in medical image analysis. However, their application in gastrointestinal endoscopy is currently hindered by two critical limitations: the misalignment between…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Huan Zheng , Yucheng Zhou , Tianyi Yan , Dubing Chen , Hongbo Lu , Wenlong Liao , Tao He , Pai Peng , Jianbing Shen

The success of Large Language Models (LLMs) has led to a parallel rise in the development of Large Multimodal Models (LMMs), which have begun to transform a variety of applications. These sophisticated multimodal models are designed to…

Artificial Intelligence · Computer Science 2025-05-20 Fouad Trad , Ali Chehab

As advances in large language models (LLMs) and multimodal techniques continue to mature, the development of general-purpose multimodal large language models (MLLMs) has surged, offering significant applications in interpreting natural…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Yuxuan Sun , Chenglu Zhu , Sunyi Zheng , Kai Zhang , Lin Sun , Zhongyi Shui , Yunlong Zhang , Honglin Li , Lin Yang

While Multimodal Large Language Models (MLLMs) have experienced significant advancement in visual understanding and reasoning, their potential to serve as powerful, flexible, interpretable, and text-driven models for Image Quality…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Tianhe Wu , Kede Ma , Jie Liang , Yujiu Yang , Lei Zhang

Introducing interpretability and reasoning into Multiple Instance Learning (MIL) methods for Whole Slide Image (WSI) analysis is challenging, given the complexity of gigapixel slides. Traditionally, MIL interpretability is limited to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Saarthak Kapse , Pushpak Pati , Srijan Das , Jingwei Zhang , Chao Chen , Maria Vakalopoulou , Joel Saltz , Dimitris Samaras , Rajarsi R. Gupta , Prateek Prasanna