Related Experiment Videos
Development and External Validation of a Machine Learning-Based Simplified Injury Severity Score for Improved Trauma
Michael D Cobler-Lichter1, Ana M Reyes1, Jessica M Delamater1
1. Division of Trauma & Surgical Critical Care, DeWitt Daughtry Family Department of Surgery, Ryder Trauma Center, University of Miami Miller School of Medicine, Miami, FL.
Background:
Trauma quality improvement requires timely, accurate injury severity quantification. Injury Severity Score (ISS), the standard measure, is typically calculated after discharge, requires manual abstraction of Abbreviated Injury Scale scores, and excludes physiologic information. We hypothesized that a machine learning (ML)-based severity score using early clinical variables would improve mortality prediction, particularly in severely injured patients.
Study Design:
ML model development used 4,521,790 patients from the American College of Surgeons Trauma Quality Improvement Program (2017-2021), with external validation in 12,858 patients from an urban Level 1 trauma center registry (2022-2025). In-hospital mortality was the primary endpoint. Multiple model architectures were compared, the best-performing pipeline was locked a priori, and discrimination versus ISS and Trauma and Injury Severity Score (TRISS) was assessed by paired bootstrap testing.
Results:
In external validation, the gradient-boosted decision tree-based ML-ISS demonstrated greater discrimination than ISS (AUROC 0.975 [95% CI 0.969-0.980] vs 0.871 [95% CI 0.852-0.888], p<0.0001) and similar discrimination to TRISS (0.973 [95% CI 0.966-0.979], p=0.239). Among patients with ISS≥30, ML-ISS outperformed both ISS (AUROC 0.913 [95% CI 0.893-0.933] vs 0.637 [95% CI 0.595-0.676], p<0.0001) and TRISS (0.897 [95% CI 0.873-0.919], p=0.021).
Conclusions:
ML-ISS improved mortality discrimination over ISS and, among severely injured patients, outperformed both ISS and TRISS. These externally validated findings support prospective multicenter evaluation of ML-ISS for trauma benchmarking and quality-improvement workflows.