Related Experiment Videos
Development and Validation of Machine Learning Models for Postjudgment Estimation of High Compensation Ratios After
Lixuan Song1, Siwen Zhao1, Yuanzheng Deng1
1Yunnan University of Chinese Medicine, Kunming, Yunnan, China.
Background:
Orthopedic surgery is the second most common subspecialty involved in medical malpractice claims, wherein lower limb surgery carries a higher risk of claims and involves higher compensation amounts. However, effective tools for postjudgment estimation of high compensation ratios and consistency assessment against prior similar cases after lower limb fracture surgery are currently lacking in medicolegal risk management and judicial practice.
Objective:
This study aimed to develop and validate multiple machine learning (ML) models to estimate a medical malpractice compensation ratio of ≥50% in postjudgment, nonfinalized medicolegal cases, and to systematically evaluate the models' discriminative ability, stability, calibration performance, and medicolegal utility.
Methods:
This study developed binary classification models based on 451 medical malpractice cases after lower limb fracture surgery in China from 2004 to 2025. Among these cases, 360 cases from eastern, northeastern, and central China constituted the development dataset and were randomly split at a 7:3 ratio into a training set (n=251) and an internal test set (n=109), whereas 91 cases from western China were held out as an independent geographical external validation set. Least Absolute Shrinkage and Selection Operator (LASSO) was used for feature selection. Logistic regression (LR), k-nearest neighbors (KNN), support vector machine (SVM), random forest (RF), and extreme gradient boosting (XGBoost) were trained using 5-fold cross-validation and grid search. Model performance was assessed using the receiver operating characteristic (ROC) curve, area under the receiver operating characteristic curve (AUROC), bootstrap resampling, calibration curves, and decision curve analysis (DCA).
Results:
LASSO identified following 8 predictors: inappropriate surgical procedure, age, inadequate medical records, lack of informed consent, disability severity grade 1-4, sex, treatment delay, and inadequate preoperative preparation. In the test set, the LR model achieved an AUROC of 0.933, with recall, precision, accuracy, and F1-score of 0.810, 0.940, 0.872, and 0.870, respectively. Inappropriate surgical procedure and lack of informed consent were the strongest model-associated factors. Bootstrap analysis showed stable LR discrimination, and calibration was favorable, with a Brier score of 0.105. DCA suggested a potential reference value across threshold probabilities. Considering discrimination, calibration, classification performance, parsimony, and interpretability, the LR model was selected as the final model. External validation provided preliminary support for cross-regional transportability.
Conclusions:
ML-based models for postjudgment estimation may provide a nonbinding historical benchmark for estimating whether a compensation ratio of ≥50% is broadly consistent with previous similar cases in medical malpractice claims. Notably, the LR model showed the most favorable overall performance among the compared models in terms of discriminative ability, stability, calibration, and medicolegal utility, providing an auxiliary postjudgment reference for hospital risk management and legal departments, legal practitioners, courts, and other judicial professionals before a case becomes final or is practically concluded.
Trial Registration:
OSF Registries 10.17605/OSF.IO/EMAKT; https://osf.io/emakt/overview.