Related Experiment Videos
Comparison of receiver operating curves derived from the same population: a bootstrapping approach
Summary
This study introduces bootstrapping for receiver operating characteristic (ROC) curves to statistically compare prediction models. It enhances the assessment of model performance differences, especially when ROC curves diverge partially.
Area of Science:
- Biostatistics
- Medical Informatics
- Epidemiology
Background:
- Receiver Operating Characteristic (ROC) curves are essential for evaluating prediction model sensitivity and specificity across decision rule cutpoints.
- Comparing two models within the same patient cohort requires robust statistical methods for their ROC curves.
- Existing methods, like Hanley et al.'s, address overall ROC curve comparison, but often ROC curves exhibit partial differences.
Purpose of the Study:
- To propose bootstrapping of ROC curves as a method for graphical assessment.
- To statistically validate differences confined to specific portions of ROC curves.
- To provide a new approach for comparing prediction models where ROC curves may cross or diverge partially.
Main Methods:
- Application of bootstrapping techniques to ROC curves derived from prediction models.
- Graphical analysis of bootstrapped ROC curves to identify significant differences.
- Statistical comparison of model performance using a case study on coronary artery disease progression prediction.
Main Results:
- Bootstrapping provides a visual and statistical tool to assess localized differences between ROC curves.
- The method is effective in detecting statistically significant variations in model performance within specific operating ranges.
- Demonstrated utility in comparing coronary artery disease prediction models, highlighting nuanced performance distinctions.
Conclusions:
- Bootstrapping ROC curves is a valuable technique for detailed statistical comparison of prediction models.
- This approach complements existing methods by offering insights into partial curve differences.
- Enhances the evaluation of model accuracy and clinical utility, particularly in complex diagnostic scenarios.