Related Experiment Video
Updated: Jun 29, 2025

09:41
A Pipeline for 3D Multimodality Image Integration and Computer-assisted Planning in Epilepsy Surgery
Published on: May 20, 2016
12.3K
Multimodal semi-supervised learning for online recognition of multi-granularity surgical workflows
Yutaro Yamada1, Jacinto Colan2, Ana Davila3
1Department of Micro-Nano Mechanical Science and Engineering, Nagoya University, Furo-cho, Chikusa-ku, Nagoya, Aichi, 464-8603, Japan. yamada@robo.mein.nagoya-u.ac.jp.
Summary
This study introduces a new semi-supervised learning method for surgical workflow recognition using multimodal data. The approach effectively learns representations from video and kinematic data, improving accuracy and reducing annotation needs.
Area of Science:
- Computer Science
- Medical Informatics
- Robotics
Background:
- Surgical workflow recognition is crucial for operating room efficiency and safety.
- Current methods often require extensive labeled data and focus on single tasks or modalities.
- Limitations include high annotation costs and lack of generalizability across diverse surgical procedures.
Purpose of the Study:
- To develop a novel semi-supervised learning approach for surgical workflow recognition.
- To leverage multimodal data (video and kinematics) and self-supervision for robust representation learning.
- To improve the efficiency of annotation while maintaining high performance in recognizing surgical gestures, phases, and steps.
Main Methods:
- A two-stage representation learning process was employed.
- Stage 1: Time contrastive learning for unsupervised spatiotemporal visual feature extraction from video.
- Stage 2: Multimodal Variational Autoencoder (VAE) fusion of visual and kinematic features, followed by recurrent neural networks for online recognition.
Main Results:
- The proposed method achieved performance comparable or superior to fully supervised models on JIGSAWS and MISAW datasets.
- Gesture recognition accuracy reached 83.3% on the JIGSAWS Suturing dataset.
- The model maintained high performance with significantly reduced annotation requirements (half the labels), demonstrating enhanced annotation efficiency.
Conclusions:
- The developed multimodal representation is versatile and effective across various surgical tasks.
- This approach significantly improves annotation efficiency for surgical workflow recognition models.
- The findings have substantial implications for advancing real-time decision-making systems in surgical environments.

