Related Experiment Video
Updated: Jun 28, 2025

04:48
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
394
A Multimodal Transformer Model for Recognition of Images from Complex Laparoscopic Surgical Videos
Rahib H Abiyev1, Mohamad Ziad Altabel1, Manal Darwish1
1Applied Artificial Intelligence Research Centre, Department of Computer Engineering, Near East University, 99132 North Cyprus, Turkey.
Diagnostics (Basel, Switzerland)
|April 13, 2024
Summary
This study introduces a multimodal artificial intelligence (AI) model to improve surgical safety by analyzing video and text data. The AI model achieved 91% accuracy in identifying critical events during laparoscopic surgery.
Area of Science:
- Surgical Safety and Artificial Intelligence
- Medical Imaging and Machine Learning
- Multimodal AI in Healthcare
Background:
- The role of artificial intelligence (AI) in surgery is not yet fully understood.
- Enhancing patient safety and reducing adverse events are critical goals in surgical practice.
- Multimodal AI offers potential for deeper insights by integrating diverse data streams.
Purpose of the Study:
- To develop and evaluate a novel multimodal AI model for surgical video analysis.
- To assess the efficacy of integrating text and image features for enhanced surgical event detection.
- To improve patient safety through AI-driven analysis of surgical procedures.
Main Methods:
- Developed a multimodal model inspired by the Video-Audio-Text Transformer architecture.
- Utilized state-of-the-art models (Vision Transformer and BERT) for text and image embedding.
- Employed convolution-free Transformer architectures for feature extraction and a joint space for modality fusion.
- Trained and tested the model on laparoscopic cholecystectomy (LC) video data from the Cholec80 dataset.
Main Results:
- The model demonstrated high performance in extracting distinct features from surgical videos.
- A joint space effectively combined text and image features, preserving inter-modal relationships.
- Achieved a mean accuracy of 91.0%, precision of 81%, and recall of 83% on the test set.
- The model showed efficacy in analyzing laparoscopic cholecystectomy videos of varying complexity.
Conclusions:
- The developed multimodal AI model shows promise in enhancing surgical safety.
- Integrating visual and textual data through advanced Transformer architectures is effective.
- This approach represents a significant step towards AI-assisted surgical monitoring and risk reduction.

