English

Dual Invariance Self-training for Reliable Semi-supervised Surgical Phase Recognition

Image and Video Processing 2025-01-30 v1 Computer Vision and Pattern Recognition

Abstract

Accurate surgical phase recognition is crucial for advancing computer-assisted interventions, yet the scarcity of labeled data hinders training reliable deep learning models. Semi-supervised learning (SSL), particularly with pseudo-labeling, shows promise over fully supervised methods but often lacks reliable pseudo-label assessment mechanisms. To address this gap, we propose a novel SSL framework, Dual Invariance Self-Training (DIST), that incorporates both Temporal and Transformation Invariance to enhance surgical phase recognition. Our two-step self-training process dynamically selects reliable pseudo-labels, ensuring robust pseudo-supervision. Our approach mitigates the risk of noisy pseudo-labels, steering decision boundaries toward true data distribution and improving generalization to unseen data. Evaluations on Cataract and Cholec80 datasets show our method outperforms state-of-the-art SSL approaches, consistently surpassing both supervised and SSL baselines across various network architectures.

Keywords

Cite

@article{arxiv.2501.17628,
  title  = {Dual Invariance Self-training for Reliable Semi-supervised Surgical Phase Recognition},
  author = {Sahar Nasirihaghighi and Negin Ghamsarian and Raphael Sznitman and Klaus Schoeffmann},
  journal= {arXiv preprint arXiv:2501.17628},
  year   = {2025}
}
R2 v1 2026-06-28T21:23:45.691Z