Related Experiment Video
Updated: Aug 5, 2026

05:12
Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery
Published on: August 12, 2021
Surgical video workflow analysis via visual-language learning
Pengpeng Li1, Xiangbo Shu2, Chun-Mei Feng3
1School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing, China.
Npj Health Systems
|July 29, 2026
Summary
This study introduces a new deep learning framework for detailed surgical video analysis, accurately identifying instrument-verb-target triplets. This enhances surgical scene understanding and decision-making in computer-assisted surgery.
Area of Science:
- Computer-assisted surgery
- Medical image analysis
- Deep learning in healthcare
Background:
- Current surgical video analysis often uses coarse-grained methods like phase or instrument recognition.
- Existing triplet recognition models have limited scope, focusing only on intra-triplet relationships.
Purpose of the Study:
- To develop a fine-grained analysis method for surgical videos by accurately identifying
triplets. - To improve surgical scene understanding and decision-making in computer-assisted surgery.
Main Methods:
- Proposed a vision-language deep learning framework (I²TM) incorporating intra- and inter-triplet modeling.
- Developed a surgical triplet semantic enhancer (TSE) for cross-modal semantic relationship establishment.
Main Results:
- The I²TM framework effectively captures finer semantics in surgical videos.
- Demonstrated enhanced accuracy and robustness in surgical triplet recognition on benchmark datasets.
Conclusions:
- The proposed approach enables comprehensive fine-grained surgical video understanding and analysis.
- This method holds significant potential for widespread applications in medical fields.