中文
相关论文

相关论文: TC-SSA: Token Compression via Semantic Slot Aggreg…

200 篇论文

In digital pathology, Whole Slide Image (WSI) analysis is usually formulated as a Multiple Instance Learning (MIL) problem. Although transformer-based architectures have been used for WSI classification, these methods require modifications…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Juan I. Pisula , Katarzyna Bozek

Unsupervised domain adaptation for medical image segmentation remains a significant challenge due to substantial domain shifts across imaging modalities, such as CT and MRI. While recent vision-language representation learning methods have…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Lalit Maurya , Honghai Liu , Reyer Zwiggelaar

Multimodal survival methods combining gigapixel histology whole-slide images (WSIs) and transcriptomic profiles are particularly promising for patient prognostication and stratification. Current approaches involve tokenizing the WSIs into…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Andrew H. Song , Richard J. Chen , Guillaume Jaume , Anurag J. Vaidya , Alexander S. Baras , Faisal Mahmood

Standard language models employ unique, monolithic embeddings for each token, potentially limiting their ability to capture the multifaceted nature of word meanings. We investigate whether tokens can be more effectively represented through…

计算与语言 · 计算机科学 2025-09-24 Kavin R , Pawan Goyal

Vision-Language-Action (VLA) models have shown remarkable promise in robotics manipulation, yet their high computational cost hinders real-time deployment. Existing token pruning methods suffer from a fundamental trade-off: aggressive…

机器人学 · 计算机科学 2026-05-19 Yixu Feng , Zinan Zhao , Yanxiang Ma , Chenghao Xia , Chengbin Du , Yunke Wang , Chang Xu

Presenting whole slide images (WSIs) as graph will enable a more efficient and accurate learning framework for cancer diagnosis. Due to the fact that a single WSI consists of billions of pixels and there is a lack of vast annotated datasets…

图像与视频处理 · 电气工程与系统科学 2023-06-09 Milan Aryal , Nasim Yahyasoltani

In medical image segmentation tasks, the domain gap caused by the difference in data collection between training and testing data seriously hinders the deployment of pre-trained models in clinical practice. Continual Test-Time Adaptation…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Xiaogang Du , Jiawei Zhang , Tongfei Liu , Tao Lei , Yingbo Wang

Existing WSI analysis methods lie on the consensus that histopathological characteristics of tumors are significant guidance for cancer diagnostics. Particularly, as the evolution of cancers is a continuous process, the correlations and…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Tong Shu , Jun Shi , Dongdong Sun , Zhiguo Jiang , Yushan Zheng

Cropping high-resolution document images into multiple sub-images is the most widely used approach for current Multimodal Large Language Models (MLLMs) to do document understanding. Most of current document understanding methods preserve…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Renshan Zhang , Yibo Lyu , Rui Shao , Gongwei Chen , Weili Guan , Liqiang Nie

Multiple instance learning (MIL) has become a preferred method for gigapixel whole slide image (WSI) classification without requiring patch-level annotations. Current MIL research primarily relies on embedding-based approaches, which…

计算机视觉与模式识别 · 计算机科学 2025-03-10 Bryan Wong , Sungrae Hong , Mun Yong Yi

To manage the complexity of transformers in video compression, local attention mechanisms are a practical necessity. The common approach of partitioning frames into patches, however, creates architectural flaws like irregular receptive…

图像与视频处理 · 电气工程与系统科学 2025-10-07 Alexander Kopte , André Kaup

State-of-the-art (SOTA) gradient-based adversarial attacks on spiking neural networks (SNNs), which largely rely on extending FGSM and PGD frameworks, face a critical limitation: substantial attack latency from multi-timestep processing,…

计算机视觉与模式识别 · 计算机科学 2025-08-20 Donghwa Kang , Doohyun Kim , Sang-Ki Ko , Jinkyu Lee , Hyeongboo Baek , Brent ByungHoon Kang

Foundation models are reshaping computational histopathology, yet their value for whole-slide image retrieval relative to strong patch-based and supervised aggregation baselines remains unclear. We benchmarked ten pipelines on 9,387…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Tianhao Lei , Parsa Esmaeilkhani , Saghir Alfasly , Wataru Uegami , Judy C. Boughey , Matthew P. Goetz , Krishna R. Kalari , H. R. Tizhoosh

Whole slide imaging is routinely adopted for carcinoma diagnosis and prognosis. Abundant experience is required for pathologists to achieve accurate and reliable diagnostic results of whole slide images (WSI). The huge size and…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Pingyi Chen , Chenglu Zhu , Sunyi Zheng , Honglin Li , Lin Yang

In this paper, we present LLaVA-Scissor, a training-free token compression strategy designed for video multimodal large language models. Previous methods mostly attempt to compress tokens based on attention scores, but fail to effectively…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Boyuan Sun , Jiaxing Zhao , Xihan Wei , Qibin Hou

Multiple instance learning (MIL) has become the leading approach for extracting discriminative features from whole slide images (WSIs) in computational pathology. Attention-based MIL methods can identify key patches but tend to overlook…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Lubin Gan , Xiaoman Wu , Jing Zhang , Zhifeng Wang , Linhao Qu , Siying Wu , Xiaoyan Sun

Deep learning is a powerful tool for whole slide image (WSI) analysis. Typically, when performing supervised deep learning, a WSI is divided into small patches, trained and the outcomes are aggregated to estimate disease grade. However,…

计算机视觉与模式识别 · 计算机科学 2022-05-20 Yi Zheng , Rushin H. Gindra , Emily J. Green , Eric J. Burks , Margrit Betke , Jennifer E. Beane , Vijaya B. Kolachalama

Accurate diagnosis of disease often depends on the exhaustive examination of Whole Slide Images (WSI) at microscopic resolution. Efficient handling of these data-intensive images requires lossy compression techniques. This paper…

Large Language Models (LLMs) have achieved remarkable performance across a wide range of Natural Language Processing (NLP) tasks. However, in long-context scenarios, they face two challenges: high computational cost and information…

The multi-scale information among the whole slide images (WSIs) is essential for cancer diagnosis. Although the existing multi-scale vision Transformer has shown its effectiveness for learning multi-scale image representation, it still…

计算机视觉与模式识别 · 计算机科学 2023-05-26 Saisai Ding , Juncheng Li , Jun Wang , Shihui Ying , Jun Shi