Related Experiment Video
Updated: Sep 19, 2025

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
3.0K
Deep learning approaches to surgical video segmentation and object detection: A scoping review
Devanish N Kamtam1, Joseph B Shrager2, Satya Deepya Malla1
1Division of Thoracic Surgery, Department of Cardiothoracic Surgery, Stanford University School of Medicine, Stanford, CA, USA.
Computers in Biology and Medicine
|June 3, 2025
Summary
Deep learning-based computer vision models show real-time potential for segmenting anatomical structures in surgical videos, especially larger organs. Further research is needed for smaller structures and generalizability.
Area of Science:
- Biomedical engineering
- Computer vision
- Surgical technology
Background:
- Computer vision (CV) has revolutionized medical imaging fields like radiology and pathology.
- Its application in real-time surgical procedures remains limited, despite its potential.
- This review focuses on deep learning (DL) for anatomical structure segmentation in surgical videos.
Purpose of the Study:
- To evaluate the state-of-the-art performance of DL-based CV models for semantic segmentation and object detection in surgical videos.
- To examine the progress of these models toward clinical applications.
- To identify challenges in segmenting surgical anatomical structures.
Main Methods:
- A scoping review of studies from 2014-2024 was conducted.
- Searches were performed in PubMed, Embase, and IEEE Xplore.
- The review focused on semantic segmentation and object detection of anatomical structures.
Main Results:
- 61 studies were identified, primarily focusing on general, colorectal, and neurosurgery.
- Semantic segmentation was the main CV task, with U-Net and DeepLab being common models.
- Model performance varied by structure size, with larger organs (e.g., liver) achieving higher accuracy (Dice score: 0.88) than smaller structures (e.g., nerves, Dice score: 0.49).
- Real-time inference speeds ranged from 5 to 298 frames per second.
Conclusions:
- Significant progress in DL-based semantic segmentation for surgical videos demonstrates real-time applicability, particularly for larger organs.
- Future advancements require addressing challenges related to smaller structures, data availability, and model generalizability.

