Integrating Clinical Priorities into the Technical Validation of Machine Learning Classifiers: A Generalizable
1Laboratory of Algorithmic Medicine, Department of Osteopathic Manipulative Medicine, College of Osteopathic Medicine, New York Institute of Technology, Old Westbury, NY 11568, USA.
Abstract:
Background/Objectives: Traditional machine learning metrics often fail to capture the clinical consequences of classification errors, particularly in high-stakes screening applications. When applying task-specific AI tools to human lives, clinicians cannot rely on simple, headline metrics and marketing materials. Impressive accuracy figures can hide dangerous algorithmic shortcuts that collapse when encountering actual patients. This study addresses the gap between statistical performance and clinical utility by evaluating classifiers for prescription opioid misuse using a clinically oriented framework. Methods: Three ensemble models, namely, Cost-Sensitive Bagged Trees (CSBT-Untuned and CSBT-Tuned) and Random Undersampling Boosting (RUSBoost), were trained on National Survey on Drug Use and Health data and assessed using composite metrics integrating clinical priorities and asymmetric error costs. Results: The results demonstrate that while CSBT-Untuned achieved the highest raw accuracy of 86.7%, it missed 61.5% of positive cases. Conversely, the threshold-optimized CSBT-Tuned model achieved enhanced minority class detection and numerically higher Clinical Discriminative Performance Scores under sensitivity-priority scenarios, though its performance remained statistically comparable to RUSBoost given overlapping confidence intervals. Learning curve analysis confirmed stable convergence for the cost-sensitive bagging approach. Conclusions: Before accepting AI tools in patient care, physicians must demand an evaluation report similarly detailed to this manuscript, utilizing this framework as guidance for what an evidence report should demonstrate prior to clinical deployment.
