English

End to End AI System for Surgical Gesture Sequence Recognition and Clinical Outcome Prediction

Artificial Intelligence 2025-11-18 v1 Computer Vision and Pattern Recognition

Abstract

Fine-grained analysis of intraoperative behavior and its impact on patient outcomes remain a longstanding challenge. We present Frame-to-Outcome (F2O), an end-to-end system that translates tissue dissection videos into gesture sequences and uncovers patterns associated with postoperative outcomes. Leveraging transformer-based spatial and temporal modeling and frame-wise classification, F2O robustly detects consecutive short (~2 seconds) gestures in the nerve-sparing step of robot-assisted radical prostatectomy (AUC: 0.80 frame-level; 0.81 video-level). F2O-derived features (gesture frequency, duration, and transitions) predicted postoperative outcomes with accuracy comparable to human annotations (0.79 vs. 0.75; overlapping 95% CI). Across 25 shared features, effect size directions were concordant with small differences (~ 0.07), and strong correlation (r = 0.96, p < 1e-14). F2O also captured key patterns linked to erectile function recovery, including prolonged tissue peeling and reduced energy use. By enabling automatic interpretable assessment, F2O establishes a foundation for data-driven surgical feedback and prospective clinical decision support.

Keywords

Cite

@article{arxiv.2511.11899,
  title  = {End to End AI System for Surgical Gesture Sequence Recognition and Clinical Outcome Prediction},
  author = {Xi Li and Nicholas Matsumoto and Ujjwal Pasupulety and Atharva Deo and Cherine Yang and Jay Moran and Miguel E. Hernandez and Peter Wager and Jasmine Lin and Jeanine Kim and Alvin C. Goh and Christian Wagner and Geoffrey A. Sonn and Andrew J. Hung},
  journal= {arXiv preprint arXiv:2511.11899},
  year   = {2025}
}