Related Experiment Video
Updated: Aug 16, 2025

07:46
Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility
Published on: August 9, 2024
822
Surgical Gesture Recognition in Laparoscopic Tasks Based on the Transformer Network and Self-Supervised Learning
Athanasios Gazis1, Pantelis Karaiskos1, Constantinos Loukas1
1Laboratory of Medical Physics, Medical School, National and Kapodistrian University of Athens, 115 27 Athens, Greece.
Bioengineering (Basel, Switzerland)
|December 23, 2022
Summary
This study introduces a deep learning framework for surgical gesture recognition, achieving high accuracy in laparoscopic tasks. A self-supervised model demonstrates comparable performance to supervised methods, highlighting its potential for efficient training.
Area of Science:
- Computer Vision
- Medical Robotics
- Machine Learning
Background:
- Accurate surgical gesture recognition is crucial for improving surgical training and performance.
- Video-based analysis offers a non-invasive method for assessing surgical skills.
Purpose of the Study:
- To develop and evaluate a deep learning framework for video-based surgical gesture recognition.
- To investigate the effectiveness of a self-supervision scheme for training surgical gesture recognition models.
Main Methods:
- A modular framework combining 3D Convolutional Networks (CNNs) for spatial-temporal features and Transformer networks for long-term dependencies.
- Two models were proposed: C3DTrans (supervised) and SSC3DTrans (self-supervised).
- Models were trained and evaluated on laparoscopic peg transfer and knot tying datasets, including the JIGSAWS robotic surgery dataset.
Main Results:
- The supervised C3DTrans model achieved high accuracy (88.0% overall, up to 97.9% gesture-level).
- The self-supervised SSC3DTrans model showed comparable performance to C3DTrans when trained on 60% of the data.
- C3DTrans demonstrated competitive performance on the JIGSAWS dataset, outperforming prior single-stream methods.
Conclusions:
- The proposed deep learning framework effectively recognizes surgical gestures from video data.
- Self-supervision presents a viable alternative for training surgical gesture recognition models, especially with limited annotated data.
- The framework shows promise for real-world applications in robotic surgery and surgical training.

