Related Experiment Video
Updated: Mar 6, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Predicting the Value of Radiology Artificial Intelligence Applications: Large-Scale Predeployment Evaluation of a
David B Larson1, Jason A Poff2, Sriyesh Krishnan2
1Department of Radiology, AI Development and Evaluation (AIDE) Lab, Stanford University School of Medicine, 453 Quarry Rd, Mail Code 5659, Stanford, CA 94304.
None:
BACKGROUND. Real-world performance of radiology artificial intelligence (AI) applications frequently diverges from previously reported results, creating challenges in anticipating a model's clinical value and impact. OBJECTIVE. The purpose of this study was to develop a structured predeployment evaluation method for radiology AI models that combines standard performance metrics with new augmentation metrics in predicting the overall value of an AI model and to test this method's predictions against radiologists' real-world postdeployment perceptions of model value. METHODS. In this prospective study, from July 2022 to November 2024, a large national radiology practice conducted a predeployment evaluation of a single vendor's portfolio of 13 AI models for 12 clinical tasks. A four-radiologist workgroup identified attributes contributing to the inherent value of AI assistance for clinical tasks, assigned weights to those attributes, and rated models accordingly. Performance of radiologists (based on clinical reports) and AI was assessed for 88,645 examinations across clinical sites by use of conventional metrics and augmentation metrics reflecting enhanced-detection cases (i.e., AI-detected radiologist-missed positive cases). The workgroup combined inherent task values and pooled AI performance to predict models' overall value. Radiologists completed a postdeployment survey. RESULTS. The workgroup identified three attributes most likely to contribute to the inherent value of AI assistance: the tediousness of the task, the likelihood that the radiologist would miss the finding, and the potential clinical impact of a missed finding. Five, five, and two tasks were rated as having high, medium, and low inherent value, respectively. Across tasks, radiologists generally had higher PPV, whereas AI generally had higher sensitivity. Models showed widely varying absolute and relative enhanced detection rates (0.03-2.27% and 4.5-60.5%, respectively). Five, five, and three models were predicted to have high, medium, and low overall value, respectively. The survey response rate was 43.2% (54/125). Perceived value categories agreed between survey respondents and workgroup predictions for 10 of 12 tasks. CONCLUSION. We present a structured method for predeployment evaluation of AI models' potential value, combining task-inherent value assessments with radiologist and AI performance metrics. A validation survey indicated high agreement between predeployment predictions and real-world postdeployment value perceptions. CLINICAL IMPACT. This practical evaluation approach can help guide radiology practices in evidence-based purchasing and deployment decisions for radiology AI models.