Related Experiment Video
Updated: Jun 29, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
SurgLaVi: Large-scale hierarchical dataset for surgical vision-language representation learning
Alejandra Perez1, Chinedu Nwoye2, Ramtin Raji Kermani2
1Intuitive Surgical, Inc., Sunnyvale, CA, United States; Center for Research and Formation in Artificial Intelligence (CinfonIA), Universidad de los Andes, Bogotá, Colombia.
Researchers developed SurgLaVi, a large surgical vision-language dataset, to improve AI understanding of surgical videos. This dataset enables better surgical workflow analysis and AI model training for enhanced surgical care.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Medical Informatics
Background:
- Vision-language pre-training (VLP) shows promise for surgical video analysis, but is limited by small, non-diverse datasets.
- Existing surgical VLP datasets lack scale, procedural variety, semantic depth, and hierarchical structure.
Purpose of the Study:
- Introduce SurgLaVi, the largest and most diverse surgical vision-language dataset.
- Enable advanced AI models for surgical workflow understanding and generalization.
- Provide an accessible, open-source dataset (SurgLaVi-β) for broader research.
Main Methods:
- Developed a fully automated pipeline for generating fine-grained surgical video transcriptions and segmentations.
- Implemented dual-modality filtering to ensure high-quality, semantically rich annotations.
- Created SurgCLIP, a CLIP-style video-text contrastive framework, to evaluate dataset utility.
Main Results:
- SurgLaVi contains nearly 240k clip-caption pairs across over 200 procedures with hierarchical annotations.
- SurgLaVi-β offers 113k clip-caption pairs from public data, significantly larger than prior datasets.
- SurgCLIP demonstrated substantial improvements in surgical phase, step, action, and tool recognition.
Conclusions:
- Large-scale, semantically rich, and hierarchically structured datasets are crucial for developing robust surgical AI.
- SurgLaVi is a key resource for advancing surgical foundation models and AI-driven surgical applications.
- The dataset facilitates improved understanding and generalization in surgical video analysis.
Related Concept Videos
Genetic Lingo
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Introduction to Language of Pathophysiology l
Introduction to Language of Pathophysiology ll

