Vision-language models for automated video analysis and documentation in laparoscopic surgery: a proof-of-concept

Esther Helene Stueker1, Fiona R Kolbinger2,3,4, Oliver Lester Saldanha1

  • 1Else Kröner Fresenius Center for Digital Health, Dresden University of Technology, Dresden, Germany.

Summary

Vision-Language Models (VLMs) show promise for surgical documentation. GPT-4o and Gemini-1.5-pro reliably detected surgical tools and classified procedures, though grading pathology requires further development.