Related Experiment Videos
SurgFusion-Net: Diversified Adaptive Multimodal Fusion Network for Surgical Skill Assessment
IEEE Transactions on Medical Imaging
|August 10, 2026
Summary
This study introduces new datasets and a multimodal fusion network for robotic-assisted surgery skill assessment. The SurgFusion-Net improves accuracy by integrating RGB, optical flow, and segmentation masks, outperforming existing methods in clinical settings.
Area of Science:
- Robotics
- Computer Vision
- Surgical Analytics
Background:
- Robotic-assisted surgery (RAS) is established, but automated skill assessment faces challenges due to complex data, limited datasets, and fusion techniques.
- Current methods using only RGB video fail to bridge the domain gap to clinical procedures, struggling with visual challenges like tissue movement and reflections.
- Clinical RAS cases present unique difficulties, including background peristalsis and specular reflections from instruments and needles.
Purpose of the Study:
- To introduce two novel clinical datasets, RAH-skill and RARP-skill, for robotic-assisted surgery skill assessment.
- To propose SurgFusion-Net, a multimodal fusion network designed to overcome visual challenges in clinical RAS skill assessment.
- To enhance the accuracy and robustness of automated surgical skill assessment in real-world clinical settings.
Main Methods:
- Developed RAH-skill and RARP-skill datasets with RGB frames, optical flow, and segmentation masks, paired with expert M-GEARS annotations.
- Proposed SurgFusion-Net, a multimodal fusion network integrating RGB, optical flow, and tool segmentation masks.
- Introduced Divergence Regulated Attention (DRA) for adaptive cross-modal feature alignment to handle visual complexities.
Main Results:
- SurgFusion-Net demonstrated superior performance on the JIGSAWS benchmark, RAH-skill, and RARP-skill datasets.
- Achieved significant Spearman rank correlation coefficient (SCC) improvements: 0.02 (LOSO) and 0.04 (LOUO) on JIGSAWS, and 0.0538 (RAH-skill) and 0.0493 (RARP-skill) on new datasets.
- The proposed method effectively addresses background motion noise and instrument specular reflections in clinical RAS.
Conclusions:
- The developed datasets and SurgFusion-Net provide a robust solution for automated skill assessment in clinical robotic-assisted surgery.
- The multimodal approach, particularly with DRA, enhances accuracy and generalizability across diverse surgical settings.
- This work contributes to advancing surgical analytics and education through improved automated skill evaluation.