English

RADA: Region-Aware Dual-encoder Auxiliary learning for Barely-supervised Medical Image Segmentation

Computer Vision and Pattern Recognition 2026-04-14 v1

Abstract

Deep learning has greatly advanced medical image segmentation, but its success relies heavily on fully supervised learning, which requires dense annotations that are costly and time-consuming for 3D volumetric scans. Barely-supervised learning reduces annotation burden by using only a few labeled slices per volume. Existing methods typically propagate sparse annotations to unlabeled slices through geometric continuity to generate pseudo-labels, but this strategy lacks semantic understanding, often resulting in low-quality pseudo-labels. Furthermore, medical image segmentation is inherently a pixel-level visual understanding task, where accuracy fundamentally depends on the quality of local, fine-grained visual features. Inspired by this, we propose RADA, a novel Region-Aware Dual-encoder Auxiliary learning pipeline which introduces a dual-encoder framework pre-trained on Alpha-CLIP to extract fine-grained, region-specific visual features from the original images and limited annotations. The framework combines image-level fine-grained visual features with text-level semantic guidance, providing region-aware semantic supervision that bridges image-level semantics and pixel-level segmentation. Integrated into a triple-view training framework, RADA achieves SOTA performance under extremely sparse annotation settings on LA2018, KiTS19 and LiTS, demonstrating robust generalization across diverse datasets.

Keywords

Cite

@article{arxiv.2604.11164,
  title  = {RADA: Region-Aware Dual-encoder Auxiliary learning for Barely-supervised Medical Image Segmentation},
  author = {Shuang Zeng and Boxu Xie and Lei Zhu and Xinliang Zhang and Jiakui Hu and Zhengjian Yao and Yuanwei Li and Yuxing Lu and Yanye Lu},
  journal= {arXiv preprint arXiv:2604.11164},
  year   = {2026}
}
R2 v1 2026-07-01T12:05:53.211Z