Related Experiment Video
Updated: Oct 16, 2025

07:46
Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility
Published on: August 9, 2024
892
SD-Net: joint surgical gesture recognition and skill assessment.
Jinglu Zhang1, Yinyu Nie2, Yao Lyu1
1National Centre for Computer Animation, Bournemouth University, Bournemouth, UK.
International Journal of Computer Assisted Radiology and Surgery
|October 16, 2021
Summary
We developed SD-Net, a novel deep learning model for surgical gesture recognition and skill assessment using only RGB videos. This method significantly improves accuracy in recognizing surgical actions and evaluating surgeon proficiency.
Area of Science:
- Computer Vision
- Medical Robotics
- Artificial Intelligence in Surgery
Background:
- Surgical gesture recognition is crucial for intraoperative assistance and resource management.
- Existing methods struggle with long-term temporal information and often require extra sensors.
- Accurate surgical skill assessment is vital for training and quality improvement.
Purpose of the Study:
- To propose a novel deep learning architecture, SD-Net, for joint surgical gesture recognition and skill assessment.
- To overcome limitations of previous methods by utilizing only RGB surgical video sequences.
- To enhance intraoperative context-aware assistance and clinical resource scheduling.
Main Methods:
- Utilized symmetric 1D temporal dilated convolution layers to hierarchically capture gesture clues across different time spans.
- Integrated a self-attention network to compute global frame-to-frame relativity for comprehensive feature aggregation.
- Employed a multi-task learning approach for simultaneous gesture recognition and skill assessment.
Main Results:
- Achieved state-of-the-art performance on a robotic suturing task from the JIGSAWS dataset.
- Significantly outperformed existing methods in gesture recognition, improving frame-wise accuracy by up to 6% and F1@50 score by 8%.
- Maintained 100% accuracy in surgical skill assessment using a leave-one-subject-out (LOSO) validation scheme.
Conclusions:
- The proposed SD-Net effectively extracts representative surgical video features by considering spatial, temporal, and relational contexts.
- Multi-task learning demonstrated a complementary effect between surgical skill assessment and gesture recognition.
- The method shows promise for enhancing surgical training, performance feedback, and intraoperative decision-making.

