Related Experiment Videos
Physician Ratings and Adoption of AI-Generated e-Consultation Advice: A Randomized Clinical Vignette Study
Gabrielle Desjardins1, Varun Jain1, Michael G Simonson1
1Division of General Internal Medicine, Department of Medicine, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA.
Background:
Medically specialized generative artificial intelligence (GAI) systems are increasingly used by clinicians as a form of electronic consultation (e-consultation), yet little is known about how clinicians perceive the quality of GAI-generated consultation advice or how it influences the management of clinical scenarios.
Objective:
To assess physicians' ratings of GAI-generated e-consultation advice for management decisions and to examine whether labeling the e-consultation source as human or GAI influences perceived quality or uptake of recommendations.
Design:
Randomized clinical vignette study.
Participants:
Internal medicine teaching faculty recruited from a single large academic medical center.
Main Measures:
Participants reviewed four clinical vignettes and rated e-consultation advice generated by a medically specialized GAI (OpenEvidence) for each vignette on completeness, actionability, clarity, appropriateness, trustworthiness, and use of evidence-based practice on a 5-point Likert scale. Participants were randomized to view advice labeled as being written by either a "specialist attending physician" or a "medically specialized artificial intelligence system." Changes in management decisions before and after viewing e-consultation advice were assessed by multiple-choice questions.
Results:
Forty-four faculty completed the study. Mean perceived quality ratings were high across all domains (overall means 4.2-4.7 out of 5), with no significant differences between randomization groups (all p > 0.05). Participants chose the consult-recommended management choice more frequently after viewing e-consultation advice than before (81.8% vs. 37.5%; OR 8.43, 95% CI 5.00-14.28; p < 0.001). Labeling the advice as human- or GAI-generated did not significantly affect adoption of e-consult recommendations (OR 1.18, 95% CI 0.51-2.75; p = 0.70).
Conclusions:
Internal medicine teaching faculty rated GAI-generated advice highly and frequently incorporated its recommendations into clinical decision-making for vignettes regardless of whether they were told the advice was human- or GAI-generated. Future studies should explore what factors influence the adoption of AI-generated recommendations and how curricula for faculty and trainees can shape consultation utilization.