Related Experiment Video
Updated: Jul 1, 2026

Application of the En Bloc Concept Combined with Anatomic Resection in Laparoscopic Hepatectomy
Published on: March 10, 2023
Deriving Clavien-Dindo Classification from Administrative Data: Development and External Validation in Hepatobiliary
Stylianos Tzedakis1,2, Louis Romengas2,3, Diana Berzan1
1Service de chirurgie digestive, hépatobiliaire et endocrinienne, AP-HP Centre, Groupe Hospitalier Cochin Port Royal, Paris, France.
Introduction:
Routine electronic health records (HER)/administrative data could enable time- and cost-efficient, real-time surveillance and health-economic evaluation of hepatobiliary postoperative complications, yet no validated algorithm currently exists. We aimed to develop and externally validate an interpretable procedure-code-based, internationally portable algorithm classifying 30-day complications by Clavien-Dindo (CDC) grades and to compare its performance to machine learning (ML) approaches.
Method:
A retrospective cohort study was conducted across two French tertiary hepatobiliary centers (Cochin for development; Beaujon for validation) from 2021-2023. Gold-standard CDC grades (≤II, III-IV, V) were assigned by an independent hepatobiliary surgeon blinded to algorithmic rules. The algorithm was expert-derived using 311 procedure codes mapped to 168 WHO's International Classification codes (ICHI), applying a temporal rule to ICU-related codes (≥POD4 for CDC-IV). ML comparators included RandomForest, ElasticNet and XGBoost trained using repeated cross-validation with hyperparameter tuning. Primary outcome was agreement with the gold-standard CDC classification; metrics included macro-F1-score, macro-balanced accuracy (MBA), sensitivity/specificity, weighted-kappa, and Ranked Probability Score (RPS) with 2000-bootstrap 95%CIs.
Results:
Among 959 liver resections (development: 476; validation: 488), major complications occurred in 18%, with 2.6% mortality. In validation, the expert algorithm achieved macro-F1-score: 0.962 (0.946-0.977), MBA: 0.974 (0.963-0.985), sensitivity: 0.950 (0.901-0.988), specificity 0.971 (0.951-0.985), weighted-κ: 0.928 (0.900-0.962) and RPS 0.016 (0.009-0.025). ML pipelines underperformed in all metrics. Misclassification (3.2%) was mainly due to ICU timing or incomplete coding.
Conclusion:
An interpretable therapeutic-act-based algorithm accurately reproduced CDC grading from routine data and outperformed ML approaches. ICHI mapping supports international portability for real-time complication surveillance, quality benchmarking and policy evaluation.

