Related Experiment Videos
Exploring Prototype Networks for Surgical Vision: Interpretability and Performance in Semantic Segmentation and
Yiping Li1, Ronald L P D de Jong1, Franco Badaloni2
1Department of Biomedical Engineering, Eindhoven University of Technology, 5612 AP Eindhoven, The Netherlands.
Abstract:
Deep learning-based surgical vision systems achieve strong performance in semantic segmentation and phase recognition, but their black-box nature limits traceability in safety-critical clinical settings. Prototype-based networks offer an interpretable alternative by grounding predictions in learned visual exemplars, yet their suitability for surgical video understanding remains insufficiently characterized. We adapted a prototype-based architecture to two surgical datasets, laparoscopic cholecystectomy and robot-assisted minimally invasive esophagectomy (RAMIE), and benchmarked it against conventional baselines. We evaluated a segmentation-only setting, in which prototype size and capacity were ablated, and a multitask setting, in which three strategies for coupling prototype learning to semantic segmentation and surgical phase recognition were compared. Prototype-based models underperformed the conventional baselines across both tasks and datasets. In the segmentation-only setting, the selected prototype configurations achieved Dice scores of 72.23% on Cholecystectomy and 72.07% on RAMIE, compared with 74.35% and 74.02% for the corresponding conventional baselines, and showed weaker boundary agreement. In the multitask setting, the best prototype strategy recovered competitive segmentation performance but remained 6-12 F1 points below the conventional baseline for phase recognition. Qualitatively, prototype activation maps exposed intra-structure decompositions and contextual cues that are not directly available from black-box baselines. Prototype networks provide spatially traceable evidence for surgical scene understanding, but currently trade interpretability for reduced boundary precision and phase-recognition performance. These findings motivate future work on scene-level and temporally aware prototypes for explainable surgical AI.