EndoChat: Grounded multimodal large language model for endoscopic surgery

Guankun Wang1, Long Bai2, Junyi Wang3

  • 1The Chinese University of Hong Kong, 999077, Hong Kong Special Administrative Region of China; Theory Lab, Central Research Institute, 2012 Labs, Huawei Technologies Co. Ltd., 999077, Hong Kong Special Administrative Region of China.

Medical Image Analysis
|September 10, 2025
PubMed
Summary

EndoChat, a new Multimodal Large Language Model (MLLM), enhances robotic-assisted surgery training and decision-making. It excels in endoscopic procedure understanding, offering advanced surgical scene analysis and dialogue capabilities.