Related Experiment Videos
Artificial intelligence-based diagnostic decision support for primary care of older adults using collective
Christopher D Streiffer1,2, Matthew J Press2,3, Nicholas S Bishop1
1Palliative and Advanced Illness Research (PAIR) Center, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA 19104, United States.
Objective:
Older adults face disproportionate risk of diagnostic errors in primary care, yet effective diagnostic decision support systems (DDSSs) remain lacking. Existing artificial intelligence (AI)-based DDSSs rely on gold-standard labels infeasible to obtain, have narrow diagnostic scope, and fail to reflect real-world uncertainty. We developed and validated a deep learning DDSS using imitation learning and collective clinician intelligence to generate diagnostic and order recommendations.
Materials And Methods:
We trained a multi-label neural network on structured and unstructured electronic health record data from primary care encounters of adults ≥65 years old to imitate clinician decisions. Models were evaluated using discrimination, calibration, threshold-based, and composite metrics on a temporally held-out test set. Clinical validity was assessed using a randomized, blinded equivalence study comparing model-generated recommendations with observed clinician decisions.
Results:
The study included 707 598 primary care encounters, generating recommendations across 669 diagnoses and 1000 orders. The general model demonstrated excellent performance (micro c-statistic 0.995 [95% CI 0.995-0.996], macro c-statistic 0.896 [0.888-0.897], micro F1-score 0.904 [0.903-0.905], macro F1-score 0.384 [0.381-0.385], integrated calibration index (ICI) 0.0025 [0.0024-0.0026]). Order prediction was lower (micro c-statistic 0.812 [0.811-0.814], macro c-statistic 0.786 [0.777-0.786], micro F1-score 0.305 [0.303-0.307], macro F1-score 0.0286 [0.0277-0.0294], ICI 0.0145 [0.0141-0.0150]). In clinician-validation, diagnostic disagreements were statistically equivalent between model recommendations and observed practice while order disagreements were not.
Conclusion:
Imitation learning and collective clinician intelligence provide a feasible framework for generating diagnostic and ordering recommendations that reflect real-world practice without expert-adjudicated training labels. Randomized, blinded clinician evaluation pragmatically establishes preliminary clinical validity beyond traditional retrospective metrics.