Related Experiment Videos
Beyond Area Under the Receiver Operating Characteristic Curve: Evaluating Predictive Performance Metrics Under Class
Vanessa das Graças José Ventura1, Claudio Moisés Valiense de Andrade2, Jussara Marques de Almeida2
1Medical School and University Hospital, Universidade Federal de Minas Gerais, Avenida Alfredo Balena, 110, Belo Horizonte, 30130-100, Brazil, 55 31991314221.
Area under the receiver operating characteristic curve (AUROC) is insufficient for evaluating imbalanced health care data. Class-aware metrics and learning curves are essential for accurate clinical interpretation of predictive models.
Area of Science:
- Health informatics
- Machine learning in healthcare
- Clinical decision support systems
Background:
- Imbalanced outcome distributions are common in healthcare datasets, potentially distorting predictive model performance evaluation.
- The area under the receiver operating characteristic curve (AUROC) is frequently reported but limited in reflecting clinically meaningful performance under class imbalance.
Purpose of the Study:
- To examine how metric selection influences the clinical interpretation of predictive models using imbalanced, real-world healthcare data.
- To assess the impact of rebalancing strategies on model performance and interpretation.
Main Methods:
- Retrospective cohort study of 17,018 hospitalized COVID-19 patients.
- Developed extreme gradient boosting (XGBoost) models to predict kidney replacement therapy (KRT) and mortality.
- Assessed performance using AUROC, macro-F1-score, class-specific precision/recall, calibration, decision curve analysis, and learning curves.
Main Results:
- High AUROC values (0.928 for KRT, 0.945 for mortality) masked substantially lower minority class performance (e.g., KRT recall 0.372).
- Rebalancing strategies improved minority class recall but reduced precision, with minimal AUROC impact.
- AUROC remained high despite clinically relevant shifts in error distribution between false positives and false negatives.
Conclusions:
- AUROC alone is insufficient for evaluating predictive models in imbalanced healthcare scenarios, even with rebalancing.
- Routine reporting of class-aware metrics and learning curve analysis is crucial for robust clinical evaluation.
- Avoid direct translation of models into practice without comprehensive, clinically meaningful performance assessment.
Related Concept Videos
Receiver Operating Characteristic Plot
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Kaplan-Meier Approach
Bioequivalence Data: Statistical Interpretation