Related Experiment Video
Updated: May 20, 2025

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Bridging Vision and Text: Applications and Challenges of Vision-Language Models in Urological Surgery
Wouter Bogaert1, Nicolas Carl2, Karl-Friedrich Kowalewski3
1ORSI Academy, Melle, Belgium; Department of Electronics and Information Systems, University of Ghent, Ghent, Belgium.
Abstract:
Vision-language models (VLMs) integrate visual data, such as surgical videos and medical images, with textual information for advanced artificial intelligence (AI) capabilities in surgery. This mini review highlights recent developments in the application of VLMs to surgical tasks in urology, such as answering clinical questions about surgical images, recognizing surgical instruments, identifying surgical phases, and detecting errors during procedures. Despite the potential of VLMs, significant challenges remain, particularly the limited availability of high-quality data sets. Future progress depends on overcoming these limitations, enhancing the robustness and reliability of VLMs, and creating standardized data sets. Ultimately, VLMs represent a promising advance towards integrated, multimodal AI systems capable of supporting surgeons via automated guidance, educational support, and performance evaluation. PATIENT SUMMARY: Our mini review explores new artificial intelligence (AI) tools that combine visual images and text to assist surgeons during operations. These AI tools can recognize instruments, identify surgical phases, and answer questions about surgery. Improved versions could help in making surgery safer and more efficient in the future.

