Related Experiment Videos
Systematic Bias in Comparative Evaluations of Machine Learning Versus Logistic Regression for Clinical Prediction
Szu-Han Wang1, Shih-Chieh Shao2, Ming-Ren Yang3
1Department of Plastic and Reconstructive Surgery, Fu Jen Catholic University Hospital, Fu Jen Catholic University, New Taipei City, Taiwan; Graduate Institute of Biomedical Informatics, College of Medical Science and Technology, Taipei Medical University, Taipei, Taiwan; Graduate Institute of Anatomy and Cell Biology, National Taiwan University School of Medicine, Taipei, Taiwan.
Objective:
Comparative evaluations of machine learning (ML) and logistic regression (LR) for clinical prediction frequently report ML as superior, but the methodological framework producing those comparisons has received limited scrutiny. We aimed to quantify the apparent discrimination advantage of ML over LR using trauma mortality prediction as an empirical case, and to characterise the evaluation practices that shape it.
Study Design And Setting:
Systematic review and random-effects meta-analysis combined with a meta-research appraisal of evaluation practices, using trauma mortality prediction as an empirical case. MEDLINE (Ovid), Scopus, Web of Science, and Embase were searched through 15 January 2025 for studies directly comparing any ML algorithm with LR on the same dataset (PROSPERO CRD42025636303); a sensitivity search using controlled vocabulary and a broader concept structure was performed during revision. The estimand was pre-specified as within-study AUC differences between the best-performing ML model and a single LR comparator, itself a source of bias. Risk of bias was assessed with PROBAST. A co-primary analysis was restricted to studies reporting confidence intervals for both models.
Results:
Twenty studies met review-level eligibility, of which 17 (243,324 patients) contributed to the primary quantitative synthesis. The pooled AUC difference favouring ML was 0.026 (95% CI 0.009-0.043); the co-primary estimate was 0.017 (95% CI 0.005-0.029). Heterogeneity was extreme (I2=97.9%) and the 95% prediction interval crossed zero (-0.034 to 0.086). The advantage was larger for best-of-tournament ensemble methods (0.034) than for single ML algorithms (0.005), but the interaction was not statistically significant. Four of 17 assessed studies achieved low PROBAST risk of bias, pooled estimates did not differ across risk-of-bias strata, and three studies used external or temporal validation. Four convergent evaluation practices - model-selection asymmetry, reliance on internal validation, selective reporting, and AUC-only synthesis - may jointly inflate apparent ML superiority and are not addressed by current evidence-synthesis frameworks.
Conclusion:
This study contributes to the meta-science of prediction-model evaluation: it characterises the methodological architecture that produces apparent ML superiority and translates that diagnosis into minimum standards. Current evidence does not reliably support claims that ML outperforms LR; observed differences may be preferentially overestimated under prevailing evaluation conventions. We propose six minimum methodological standards - pre-specification, fair comparator design, robust external or temporal validation, calibration and decision-analytic reporting, full transparency, and bias-aware synthesis - which may also be relevant to other clinical prediction settings.
Plain Language Summary:
Computer programs that learn patterns from data (machine learning) are often reported to predict death after injury better than traditional statistical models. We reviewed published studies that compared both approaches in the same patients. Across 17 studies the difference in ranking accuracy was small, varied widely between studies, and in a new setting could favour either approach. Part of the apparent advantage may come from how the comparisons were designed and reported rather than from the methods themselves: the best of several machine-learning models is usually compared against a single, less carefully built traditional model, almost always in the same data used to develop it. Readers should be cautious about claims of superiority unless the comparison was fair, the models were tested in a different setting, and measures beyond ranking accuracy were reported.
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
Regression Toward the Mean