Related Experiment Videos
DiagnosticXchange: an open-source framework for evaluating safety, efficiency, and diagnostic reasoning in clinical
Moran Sorka1,2, Alon Gorenshtein2,3, Hillel Abramovitch3
1Faculty of Biology, Technion-Israel Institute of Technology, Haifa 3200003, Israel.
Objective:
To develop and validate an open-source evaluation framework that assesses clinical AI diagnostic systems across multiple clinically relevant dimensions (accuracy, cost, time, invasiveness, physician effort, and safety behaviors), addressing the limitations of accuracy-only benchmarks.
Materials And Methods:
We developed DiagnosticXchange, a dynamic clinical simulation platform where AI systems interact with a simulated hospital by ordering tests, requesting imaging, and performing procedures, with each action mapped to Current Procedural Terminology (CPT) codes capturing cost, time, work relative value units, and invasiveness. We validated the framework using 8 large language models on 216 peer-reviewed cases across 19 specialties (1728 sessions). Analyses included competing-risk survival analysis, unsupervised clustering of reasoning strategies, diagnostic contribution scoring, and a pilot comparison with 14 neurologists.
Results:
Three systems achieved near-identical accuracy (93.5%-94.0%) yet differed significantly in cost (P <.001; 1.75-fold between the most and least expensive top-3 systems) and 2.1-fold in physician oversight. Competing-risk analysis showed the most efficient system solved 86% of cases within $5000, while the most resource-intensive required $10 000 for equivalent accuracy. Unsupervised clustering identified 3 reasoning strategies; behavioral profiles predicted resource consumption better than system identity (AIC: 4683 vs 5056). Safety analysis revealed premature diagnosis (up to 9.3%), noncontributory invasive procedures (8.7%-29.9%), and futile invasive procedures on failed cases (3.7%-18.1%). Test-retest analysis confirmed reproducible accuracy (84.6% concordance) but substantial process variability (cost CV: 86.6%). A pilot comparison captured multi-dimensional performance of 14 neurologists alongside AI systems.
Conclusions:
Accuracy-only evaluation is insufficient for safe clinical AI deployment. DiagnosticXchange provides an open-source, reproducible framework for multi-dimensional assessment that reveals clinically consequential differences invisible to existing benchmarks.