Related Experiment Video
Updated: Sep 11, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Federated Learning for surgical vision in appendicitis classification: Results of the FedSurg EndoVis 2024 challenge
Max Kirchner1, Hanna Hoffmann1, Alexander C Jenke1
1Department of Translational Surgical Oncology, National Center for Tumor Diseases (NCT), NCT/UCC Dresden, a partnership between DKFZ, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology, and Helmholtz-Zentrum Dresden-Rossendorf (HZDR), Dresden, Germany.
Abstract:
Developing generalizable surgical AI requires multi-institutional data, yet privacy constraints preclude direct data sharing, making Federated Learning (FL) a natural candidate. Its application to complex, spatiotemporal surgical video remains largely unbenchmarked. We present the FedSurg Challenge, the first international initiative dedicated to FL in surgical vision, as a proof-of-concept evaluation using a multi-center dataset of laparoscopic appendectomies (preliminary subset of Appendix300). Three participant submissions were evaluated on generalization to an unseen clinical center and center-specific local adaptation, alongside centralized, Swarm Learning and parameter-efficient fine-tuning baselines, and constant and chance-level reference classifiers. Our analysis identifies temporal modeling as the architectural factor most consistently associated with generalization to the unseen center, although the effect is not uniform across metrics in the controlled baseline comparison, and both parameter-efficient baselines remain more costly than a constant predictor under the ordinal metric. Classifier collapse arises both from failure of the global model to transfer under domain shift and from unconstrained fine-tuning on small, imbalanced local datasets, motivating structured personalized FL with parameter-efficient fine-tuning for center-specific adaptation. Absolute performance remains far from clinical viability: even with all data pooled centrally the task reached 26.31% F1-score on the unseen center, and no clinically acceptable threshold for intraoperative appendicitis grading has been established. Paired permutation tests over the test cases resolve only large differences, and no adaptation comparison reaches significance at this sample size. By characterizing these limitations, and the limits of what the evaluation can resolve at this scale, this work establishes a methodological reference point for privacy-preserving surgical video AI.