Related Experiment Video
Updated: Jun 25, 2026

Automatic Surgery in Transcatheter Aortic Valve Replacement Using Augmented Reality
Published on: August 9, 2024
End to end AI system for surgical gesture sequence recognition and clinical outcome prediction
Xi Li1, Nicholas Matsumoto1, Ujjwal Pasupulety2
1Department of Computational Biomedicine, Center for ArtificialIntelligence Research and Education, Cedars Sinai Medical Center, Los Angeles, CA, USA.
None:
Fine-grained analysis of intraoperative behavior and its impact on patient outcomes remains a longstanding challenge. We present Frame-to-Outcome (F2O), an end-to-end system that translates tissue dissection videos into gesture sequences and uncovers patterns associated with postoperative outcomes. Leveraging transformer-based spatial and temporal modeling and frame-wise classification, F2O robustly detects consecutive short (˜2 s) gestures in the nerve-sparing step of robot-assisted radical prostatectomy (AUC: 0.80 frame-level; 0.81 video-level). F2O-derived features-gesture frequency, duration, and transitions-predicted postoperative outcomes with accuracy comparable to human annotations (0.79 vs. 0.75; overlapping 95% CI). Across 25 shared features, effect size directions were concordant with small differences (∆davg ≈ 0.07), and strong correlation (r = 0.96, p < 1 × 10-14). F2O also captured key patterns linked to erectile function recovery, including prolonged tissue peeling and reduced energy use. By enabling automatic interpretable assessment, F2O establishes a foundation for data-driven surgical feedback and prospective clinical decision support.