Related Experiment Videos
Automated surgical phase recognition in pediatric laparoscopic fundoplication using deep learning: a retrospective
Yohei Sanmoto1,2, Luning Li3, Yudai Goto1
1Department of Pediatric Surgery, University of Tsukuba Hospital, Tsukuba, Japan.
Background:
Automated surgical phase recognition may support objective workflow analysis and surgical education but has not been sufficiently evaluated in pediatric minimally invasive surgery. This study aimed to develop and evaluate a deep learning-based phase recognition model for pediatric laparoscopic fundoplication.
Methods:
We retrospectively analyzed 38 laparoscopic fundoplication videos recorded between 2015 and 2025. Videos were annotated at 1-s intervals using a predefined 10-phase workflow. Heterogeneous nonstandard intra-abdominal intervals were excluded before model training and evaluation, resulting in a nine-phase classification framework. A phase recognition model combining a frozen Vision Transformer backbone with a bidirectional gated recurrent unit network was trained using cross-entropy and temporal smoothing losses and evaluated by case-level 5-fold cross-validation with three random seeds per fold. Overall and phase-specific classification performance, misclassification patterns, and correlations and absolute differences between predicted and reference phase durations were assessed.
Results:
Across the eight surgical phases, mean accuracy, macro-F1 score, and macro-Jaccard index were 0.812 ± 0.030, 0.805 ± 0.030, and 0.682 ± 0.042, respectively. Phase-specific mean recall ranged from 0.674 to 0.902 and was lowest for the shoulder stitch phase (0.674 ± 0.135). Predicted timelines generally reflected the overall operative workflow, with misclassifications occurring predominantly between temporally or procedurally adjacent phases. Predicted and reference phase durations showed moderate to very strong correlations across phases (Spearman's ρ, 0.469-0.937), with the lowest correlation observed for hiatal closure (ρ = 0.469); hiatal closure and wrap formation also had the largest median absolute errors (5.0 min each).
Conclusions:
Automated phase recognition in pediatric laparoscopic fundoplication videos achieved promising classification performance, although discrimination between visually similar adjacent phases and case-level phase duration estimation require further improvement. These findings establish an initial benchmark for pediatric surgical video analysis and support further investigation of its applications in operative review, surgical education, and structured performance assessment.