English
Related papers

Related papers: Multi-Resolution Pathology-Language Pre-training M…

200 papers

Vision-language models in pathology enable multimodal case retrieval and automated report generation. Many of the models developed so far, however, have been trained on pathology reports that include information which cannot be inferred…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Ruben T. Lucassen , Tijn van de Luijtgaarden , Sander P. J. Moonemans , Gerben E. Breimer , Willeke A. M. Blokx , Mitko Veta

Human expertise in chemistry and biomedicine relies on contextual molecular understanding, a capability that large language models (LLMs) can extend through fine-grained alignment between molecular structures and text. Recent multimodal…

Computation and Language · Computer Science 2025-03-10 Sumin Ha , Jun Hyeong Kim , Yinhua Piao , Sun Kim

Accurate classification of focal liver lesions is crucial for diagnosis and treatment in hepatology. However, traditional supervised deep learning models depend on large-scale annotated datasets, which are often limited in medical imaging.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Song Jian , Hu Yuchang , Wang Hui , Chen Yen-Wei

Accurate survival prediction from histopathology whole-slide images (WSIs) remains challenging due to their gigapixel resolution, strong spatial heterogeneity, and complex survival distributions. We introduce a comprehensive computational…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Ardhendu Sekhar , Vasu Soni , Keshav Aske , Shivam Madnoorkar , Pranav Jeevan , Amit Sethi

Vision-Language Pre-training (VLP) has shown the merits of analysing medical images, by leveraging the semantic congruence between medical images and their corresponding reports. It efficiently learns visual representations, which in turn…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Xiaoxuan He , Yifan Yang , Xinyang Jiang , Xufang Luo , Haoji Hu , Siyun Zhao , Dongsheng Li , Yuqing Yang , Lili Qiu

In histopathology, tissue sections are typically stained using common H&E staining or special stains (MAS, PAS, PASM, etc.) to clearly visualize specific tissue structures. The rapid advancement of deep learning offers an effective solution…

Image and Video Processing · Electrical Eng. & Systems 2025-04-23 Zizhi Chen , Xinyu Zhang , Minghao Han , Yizhou Liu , Ziyun Qian , Weifeng Zhang , Xukun Zhang , Jingwei Wei , Lihua Zhang

The rapid advancements in large language models (LLMs) have unlocked their potential for multimodal tasks, where text and visual data are processed jointly. However, applying LLMs to medical imaging, particularly for chest X-rays (CXR),…

Image and Video Processing · Electrical Eng. & Systems 2025-02-11 Nicholas Evans , Stephen Baker , Miles Reed

Pre-trained Vision-Language Models (VLMs), like CLIP, exhibit strong generalization ability to downstream tasks but struggle in few-shot scenarios. Existing prompting techniques primarily focus on global text and image representations, yet…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Xin Liu , Jiamin Wu , and Wenfei Yang , Xu Zhou , Tianzhu Zhang

Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language models (LLMs) and use end-to-end training to achieve multi-modal understanding in a unified…

Artificial Intelligence · Computer Science 2025-08-14 Zixian Guo , Ming Liu , Qilong Wang , Zhilong Ji , Jinfeng Bai , Lei Zhang , Wangmeng Zuo

High-resolution (HR) image perception presents a key bottleneck for multimodal large language models (MLLMs). While visual search offers a promising solution, existing methods struggle with the trade-off between coverage and efficiency.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Liupeng Li , Haoqian Kang , Zhenyu Lu , Jinpeng Wang , Bin Chen , Ke Chen , Yaowei Wang

Multimodal large language models (MLLMs) have achieved significant success in the general field of image processing. Their emerging task generalization and freeform conversational capabilities can greatly facilitate medical diagnostic…

Image and Video Processing · Electrical Eng. & Systems 2024-09-17 Youzhu Jin , Yichen Zhang

Large language models (LLMs) have increased interest in vision language models (VLMs), which process image-text pairs as input. Studies investigating the visual understanding ability of VLMs have been proposed, but such studies are still…

Computation and Language · Computer Science 2024-06-25 Jesse Atuhurra , Iqra Ali , Tatsuya Hiraoka , Hidetaka Kamigaito , Tomoya Iwakura , Taro Watanabe

Whole slide imaging (WSI) has transformed digital pathology by enabling computational analysis of gigapixel histopathology images. Recent foundation model advances have accelerated progress in computational pathology, facilitating joint…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Peihang Wu , Zehong Chen , Lijian Xu

Modern Vision-Language Models (VLMs) exhibit unprecedented capabilities in cross-modal semantic understanding between visual and textual modalities. Given the intrinsic need for multi-modal integration in clinical applications, VLMs have…

Image and Video Processing · Electrical Eng. & Systems 2025-06-24 Haoneng Lin , Cheng Xu , Jing Qin

In diagnostic reports, experts encode complex imaging data into clinically actionable information. They describe subtle pathological findings that are meaningful in their anatomical context. Reports follow relatively consistent structures,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Felicia Bader , Philipp Seeböck , Anastasia Bartashova , Ulrike Attenberger , Georg Langs

The problem of recognizing various types of tissues present in multi-gigapixel histology images is an important fundamental pre-requisite for downstream analysis of the tumor microenvironment in a bottom-up analysis paradigm for…

Image and Video Processing · Electrical Eng. & Systems 2021-04-02 Nima Hatami , Mohsin Bilal , Nasir Rajpoot

Medical vision-language models (VLMs) combine computer vision (CV) and natural language processing (NLP) to analyze visual and textual medical data. Our paper reviews recent advancements in developing VLMs specialized for healthcare,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Iryna Hartsock , Ghulam Rasool

Vision-language models (VLMs) have recently been integrated into multiple instance learning (MIL) frameworks to address the challenge of few-shot, weakly supervised classification of whole slide images (WSIs). A key trend involves…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Bryan Wong , Jong Woo Kim , Huazhu Fu , Mun Yong Yi

When captioning an image, people describe objects in diverse ways, such as by using different terms and/or including details that are perceptually noteworthy to them. Descriptions can be especially unique across languages and cultures.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Kyle Buettner , Jacob T. Emmerson , Adriana Kovashka

Whole-slide image analysis via the means of computational pathology often relies on processing tessellated gigapixel images with only slide-level labels available. Applying multiple instance learning-based methods or transformer models is…

Computer Vision and Pattern Recognition · Computer Science 2023-10-20 Joshua Butke , Noriaki Hashimoto , Ichiro Takeuchi , Hiroaki Miyoshi , Koichi Ohshima , Jun Sakuma
‹ Prev 1 3 4 5 6 7 10 Next ›