Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Tactile Feedback in Robot-Assisted Minimally Invasive Surgery: A Systematic Review.

The international journal of medical robotics + computer assisted surgery : MRCAS·2024
Same author

Latent Space Search-Based Adaptive Template Generation for Enhanced Object Detection in Bin-Picking Applications.

Sensors (Basel, Switzerland)·2024
Same author

Intuitive Cell Manipulation Microscope System with Haptic Device for Intracytoplasmic Sperm Injection Simplification.

Sensors (Basel, Switzerland)·2024
Same author

Endoscope Automation Framework with Hierarchical Control and Interactive Perception for Multi-Tool Tracking in Minimally Invasive Surgery.

Sensors (Basel, Switzerland)·2023
Same author

A Concurrent Framework for Constrained Inverse Kinematics of Minimally Invasive Surgical Robots.

Sensors (Basel, Switzerland)·2023
Same author

AbAdapt: an adaptive approach to predicting antibody-antigen complex structures from sequence.

Bioinformatics advances·2023

Related Experiment Video

Updated: Jun 29, 2025

A Pipeline for 3D Multimodality Image Integration and Computer-assisted Planning in Epilepsy Surgery
09:41

A Pipeline for 3D Multimodality Image Integration and Computer-assisted Planning in Epilepsy Surgery

Published on: May 20, 2016

12.3K

Multimodal semi-supervised learning for online recognition of multi-granularity surgical workflows.

Yutaro Yamada1, Jacinto Colan2, Ana Davila3

  • 1Department of Micro-Nano Mechanical Science and Engineering, Nagoya University, Furo-cho, Chikusa-ku, Nagoya, Aichi, 464-8603, Japan. yamada@robo.mein.nagoya-u.ac.jp.

International Journal of Computer Assisted Radiology and Surgery
|April 1, 2024
PubMed
Summary

This study introduces a new semi-supervised learning method for surgical workflow recognition using multimodal data. The approach effectively learns representations from video and kinematic data, improving accuracy and reducing annotation needs.

Keywords:
Multimodal learningRobotic surgerySemi-supervised learningSurgical workflow recognition

More Related Videos

Multimodal Hierarchical Imaging of Serial Sections for Finding Specific Cellular Targets within Large Volumes
11:19

Multimodal Hierarchical Imaging of Serial Sections for Finding Specific Cellular Targets within Large Volumes

Published on: March 20, 2018

10.4K
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
04:48

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography

Published on: November 30, 2022

2.7K

Related Experiment Videos

Last Updated: Jun 29, 2025

A Pipeline for 3D Multimodality Image Integration and Computer-assisted Planning in Epilepsy Surgery
09:41

A Pipeline for 3D Multimodality Image Integration and Computer-assisted Planning in Epilepsy Surgery

Published on: May 20, 2016

12.3K
Multimodal Hierarchical Imaging of Serial Sections for Finding Specific Cellular Targets within Large Volumes
11:19

Multimodal Hierarchical Imaging of Serial Sections for Finding Specific Cellular Targets within Large Volumes

Published on: March 20, 2018

10.4K
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
04:48

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography

Published on: November 30, 2022

2.7K

Area of Science:

  • Computer Science
  • Medical Informatics
  • Robotics

Background:

  • Surgical workflow recognition is crucial for operating room efficiency and safety.
  • Current methods often require extensive labeled data and focus on single tasks or modalities.
  • Limitations include high annotation costs and lack of generalizability across diverse surgical procedures.

Purpose of the Study:

  • To develop a novel semi-supervised learning approach for surgical workflow recognition.
  • To leverage multimodal data (video and kinematics) and self-supervision for robust representation learning.
  • To improve the efficiency of annotation while maintaining high performance in recognizing surgical gestures, phases, and steps.

Main Methods:

  • A two-stage representation learning process was employed.
  • Stage 1: Time contrastive learning for unsupervised spatiotemporal visual feature extraction from video.
  • Stage 2: Multimodal Variational Autoencoder (VAE) fusion of visual and kinematic features, followed by recurrent neural networks for online recognition.

Main Results:

  • The proposed method achieved performance comparable or superior to fully supervised models on JIGSAWS and MISAW datasets.
  • Gesture recognition accuracy reached 83.3% on the JIGSAWS Suturing dataset.
  • The model maintained high performance with significantly reduced annotation requirements (half the labels), demonstrating enhanced annotation efficiency.

Conclusions:

  • The developed multimodal representation is versatile and effective across various surgical tasks.
  • This approach significantly improves annotation efficiency for surgical workflow recognition models.
  • The findings have substantial implications for advancing real-time decision-making systems in surgical environments.