Related Experiment Video
Updated: Jul 12, 2025

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.0K
SDA-CLIP: surgical visual domain adaptation using video and text labels
Yuchong Li1,2, Shuangfu Jia3, Guangbi Song4
1Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China.
Quantitative Imaging in Medicine and Surgery
|October 23, 2023
Summary
This study introduces a novel domain adaptation method, SDA-CLIP, for surgical action recognition. The method significantly improves accuracy by bridging virtual reality and clinical data, outperforming existing challenge benchmarks.
Area of Science:
- Computer Vision
- Medical Robotics
- Machine Learning
Background:
- Surgical action recognition is crucial for autonomous surgery but limited by clinical data scarcity.
- Virtual reality (VR) simulations offer a scalable solution for algorithm development via domain adaptation.
- Adapting VR-trained models to clinical settings reduces data costs and protects patient privacy.
Purpose of the Study:
- To develop a domain adaptation method for cross-domain surgical action recognition.
- To leverage contrastive language-image pretraining for bridging VR and clinical surgical data.
- To enhance the accuracy of surgical action recognition in clinical settings using simulated data.
Main Methods:
- Introduced Surgical Domain Adaptation based on Contrastive Language-Image Pretraining (SDA-CLIP).
- Utilized Vision Transformer (ViT) for video embeddings and Transformer for text embeddings.
- Employed inter- and intra-modality loss functions to align embeddings across domains.
- Evaluated on the MICCAI 2020 EndoVis Challenge SurgVisDom dataset.
Main Results:
- SDA-CLIP achieved a 65.9% weighted F1-score on the hard domain adaptation task (VR data only).
- Achieved an 84.4% weighted F1-score on the soft domain adaptation task (VR and clinical-like data).
- Significantly outperformed the top-performing team in the challenge.
Conclusions:
- SDA-CLIP effectively extracts video and textual information for improved cross-domain recognition.
- The method demonstrates high performance in adapting simulated surgical data to clinical applications.
- The developed model offers a promising approach for enhancing autonomous surgical systems.

