English
Related papers

Related papers: LoC-Path: Learning to Compress for Pathology Multi…

200 papers

With the remarkable success of large language models (LLMs) in natural language understanding and generation, multimodal large language models (MLLMs) have rapidly advanced in their ability to process data across multiple modalities. While…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Jingrui Zhang , Feng Liang , Yong Zhang , Wei Wang , Runhao Zeng , Xiping Hu

Whole Slide Imaging (WSI) has become a gold standard in cancer diagnosis, inspecting multi-scale information from cellular to tissue levels. Processing an entire WSI directly is infeasible due to GPU memory constraints; thus, Multiple…

Image and Video Processing · Electrical Eng. & Systems 2026-05-08 Tianyi Zhang , Sicheng Chen , Borui Kang , Dankai Liao , Qiaochu Xue , Bochong Zhang , Fei Xia , Zeyu Liu , Yueming Jin

While long-context large language models (LLMs) exhibit remarkable document processing capabilities, their prohibitively high training costs often hinder customized applications. To mitigate this issue, we propose \textit{Sequential…

Machine Learning · Computer Science 2025-05-23 Wenhao Li , Yuxin Zhang , Gen Luo , Daohai Yu , Rongrong Ji

The rapid digitization of histopathology slides has opened up new possibilities for computational tools in clinical and research workflows. Among these, content-based slide retrieval stands out, enabling pathologists to identify…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Hongyi Wang , Zhengjie Zhu , Jiabo Ma , Fang Wang , Yue Shi , Bo Luo , Jili Wang , Qiuyu Cai , Xiuming Zhang , Yen-Wei Chen , Lanfen Lin , Hao Chen

Batch effects arising from technical variations in histopathology staining protocols, scanners, and acquisition pipelines pose a persistent challenge for computational pathology, hindering cross-batch generalization and limiting reliable…

Machine Learning · Computer Science 2026-03-02 Xiaolong Zhang , Jianwei Zhang , Selim Sevim , Emek Demir , Ece Eksi , Xubo Song

With the advancement of Large Language Model (LLM) for natural language processing, this paper presents an intriguing finding: a frozen pre-trained LLM layer can process visual tokens for medical image segmentation tasks. Specifically, we…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Fenghe Tang , Wenxin Ma , Zhiyang He , Xiaodong Tao , Zihang Jiang , S. Kevin Zhou

With the wide adoption of language models for IR -- and specifically RAG systems -- the latency of the underlying LLM becomes a crucial bottleneck, since the long contexts of retrieved passages lead large prompts and therefore, compute…

Information Retrieval · Computer Science 2026-04-06 Cornelius Kummer , Lena Jurkschat , Michael Färber , Sahar Vahdati

The escalating computational costs of Large Language Model (LLM) inference have become a critical barrier to their widespread and sustainable deployment. While existing optimization strategies are effective, they are predominantly based on…

Machine Learning · Computer Science 2025-07-02 Yilun Zhang

Multiple instance learning (MIL) has become a preferred method for gigapixel whole slide image (WSI) classification without requiring patch-level annotations. Current MIL research primarily relies on embedding-based approaches, which…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Bryan Wong , Sungrae Hong , Mun Yong Yi

Weakly supervised semantic segmentation (WSSS) in histopathology reduces pixel-level labeling by learning from image-level labels, but it is hindered by inter-class homogeneity, intra-class heterogeneity, and CAM-induced region shrinkage…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Khang Le , Anh Mai Vu , Thi Kim Trang Vo , Ha Thach , Ngoc Bui Lam Quang , Thanh-Huy Nguyen , Minh H. N. Le , Zhu Han , Chandra Mohan , Hien Van Nguyen

Recently, scaling images to high resolution has received much attention in multimodal large language models (MLLMs). Most existing practices adopt a sliding-window-style cropping strategy to adapt to resolution increase. Such a cropping…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Mingxin Huang , Yuliang Liu , Dingkang Liang , Lianwen Jin , Xiang Bai

Digital pathology has revolutionized the field by enabling the digitization of tissue samples into whole slide images (WSIs). However, the high resolution and large size of WSIs present significant challenges when it comes to applying Deep…

Computer Vision and Pattern Recognition · Computer Science 2025-07-03 Ali Mammadov , Loïc Le Folgoc , Guillaume Hocquet , Pietro Gori

Representation learning of pathology whole-slide images (WSIs) has been has primarily relied on weak supervision with Multiple Instance Learning (MIL). However, the slide representations resulting from this approach are highly tailored to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Andrew H. Song , Richard J. Chen , Tong Ding , Drew F. K. Williamson , Guillaume Jaume , Faisal Mahmood

Vision-Language Models (VLMs) face severe memory and latency bottlenecks due to high-resolution visual tokens. While current token reduction methods theoretically save FLOPs, post-hoc pruning introduces structural overhead, failing to yield…

Artificial Intelligence · Computer Science 2026-05-28 Fengze Yang , Bo Yu , Xuewen Luo , Cathy Liu , Chenxi Liu

Efficient adaption of large language models (LLMs) on edge devices is essential for applications requiring continuous and privacy-preserving adaptation and inference. However, existing tuning techniques fall short because of the high…

Advances in optical microscopy scanning have significantly contributed to computational pathology (CPath) by converting traditional histopathological slides into whole slide images (WSIs). This development enables comprehensive digital…

Image and Video Processing · Electrical Eng. & Systems 2024-11-19 Xitong Ling , Yuanyuan Lei , Jiawen Li , Junru Cheng , Wenting Huang , Tian Guan , Jian Guan , Yonghong He

Vision-Language Models (VLMs) are expensive because the LLM processes hundreds of largely redundant visual tokens. Existing token reduction methods typically exploit \textit{either} vision-encoder saliency (broad but query-agnostic)…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Dhruv Parikh , Haoyang Fan , Rajgopal Kannan , Viktor Prasanna

Commonsense reasoning often requires both textual and visual knowledge, yet Large Language Models (LLMs) trained solely on text lack visual grounding (e.g., "what color is an emperor penguin's belly?"). Visual Language Models (VLMs) perform…

Computation and Language · Computer Science 2026-04-14 Guy Yariv , Idan Schwartz , Yossi Adi , Sagie Benaim

Recent advances in artificial intelligence (AI), in particular self-supervised learning of foundation models (FMs), are revolutionizing medical imaging and computational pathology (CPath). A constant challenge in the analysis of digital…

Current cervical cytopathology whole slide image (WSI) screening primarily relies on detection-based approaches, which are limited in performance due to the expense and time-consuming annotation process. Multiple Instance Learning (MIL), a…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Jialong Huang , Gaojie Li , Shichao Kan , Jianfeng Liu , Yixiong Liang