Related Experiment Video
Updated: Jan 8, 2026

Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery
Published on: August 12, 2021
Multimodal learning for automated robotic surgical skill assessment: integrating video and kinematic data with
Ke Chen1,2,3, Yiming Zhang2,3, Ying Weng2,3
1School of Information Engineering, Wenzhou Business College, Wenzhou, China.
Background:
Surgical skill assessment relies on in-person observation by expert surgeons, which is time-consuming and subjective. Therefore, an automated objective surgical skill assessment system that could discriminate between surgeons with different levels of expertise would be helpful for surgeon training.
Materials And Methods:
In this paper, we propose a multimodal surgical skill classification framework, which contains a combination of video and kinematic classifiers. We present a 3D Convolutional Neural Network (CNN)-based video classifier and a 1D CNN-based kinematic classifier. Furthermore, we apply Class Activation Map (CAM) to provide an explainability of the model in decision-making process. We conduct the experiments on the public datasets JIGSAWS and ROSMA. JIGSAWS contains 108 trials, and it is collected from 8 surgeons with different surgical skill level. ROSMA contains 206 trials, and it is recorded from 12 subjects.
Results:
The experimental results have shown that our framework has achieved the comparable performance on the two public datasets. The proposed multimodal learning framework outperforms single-modality in most surgical tasks.
Conclusion:
This study contributes a multimodal learning approach for automatic robotic surgical skill evaluation with kinematic data and video data. It also explains the model in decision-making process using CAM.

