English

DSAGL: Dual-Stream Attention-Guided Learning for Weakly Supervised Whole Slide Image Classification

Computer Vision and Pattern Recognition 2025-06-30 v2

Abstract

Whole-slide images (WSIs) are critical for cancer diagnosis due to their ultra-high resolution and rich semantic content. However, their massive size and the limited availability of fine-grained annotations pose substantial challenges for conventional supervised learning. We propose DSAGL (Dual-Stream Attention-Guided Learning), a novel weakly supervised classification framework that combines a teacher-student architecture with a dual-stream design. DSAGL explicitly addresses instance-level ambiguity and bag-level semantic consistency by generating multi-scale attention-based pseudo labels and guiding instance-level learning. A shared lightweight encoder (VSSMamba) enables efficient long-range dependency modeling, while a fusion-attentive module (FASA) enhances focus on sparse but diagnostically relevant regions. We further introduce a hybrid loss to enforce mutual consistency between the two streams. Experiments on CIFAR-10, NCT-CRC, and TCGA-Lung datasets demonstrate that DSAGL consistently outperforms state-of-the-art MIL baselines, achieving superior discriminative performance and robustness under weak supervision.

Keywords

Cite

@article{arxiv.2505.23341,
  title  = {DSAGL: Dual-Stream Attention-Guided Learning for Weakly Supervised Whole Slide Image Classification},
  author = {Daoxi Cao and Hangbei Cheng and Yijin Li and Ruolin Zhou and Xuehan Zhang and Xinyi Li and Binwei Li and Xuancheng Gu and Jianan Zhang and Xueyu Liu and Yongfei Wu},
  journal= {arXiv preprint arXiv:2505.23341},
  year   = {2025}
}
R2 v1 2026-07-01T02:48:13.987Z