Related Experiment Video
Updated: Apr 22, 2026

E-Patient Counseling Trial E-PACO: Computer Based Education versus Nurse Counseling for Patients to Prepare for Colonoscopy
Published on: August 1, 2019
AI at the bedside: Randomised controlled trial of ChatGPT's impact on student performance in real-patient clinical
Haroon Saloojee1, Michael C Gramanie2, Rhodasi Mwali3
1School of Clinical Medicine, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa.
Background:
Generative artificial intelligence (AI) tools are entering clinical training faster than curricula and assessments can adapt. It is currently unknown whether point-of-care access to large language models (LLMs) improves clinician performance during real-time bedside assessments.
Objective:
To evaluate the effect of allowing ChatGPT use on student performance in ward-based real-patient clinical exams.
Methods:
We conducted a parallel‑group, randomised controlled trial (2:1 allocation) across four academic hospitals in a middle-income country setting. Final‑year medical students completed a 30‑minute uninterrupted patient encounter followed by a 20‑minute assessor-led evaluation. Intervention: ChatGPT (GPT‑4o) permitted during the encounter. With only minimal training, participants' point-of-care ChatGPT use reflected self-designed approaches. Control: no digital aids. Primary outcome: overall clinical performance (0-100) scored on a standard rubric (history, examination, differential, diagnosis, investigations, management, counselling). Secondary outcomes: domain sub‑scores; observer‑rated patient interaction; student experience; patient satisfaction; subsequent summative exam scores. Analyses used ANCOVA adjusting for prior academic performance and site (α = 0.05).
Results:
Seventy‑three students were analysed (ChatGPT n = 49; control n = 24). Overall performance score did not differ (66.7 ± 9.9 vs 68.2 ± 7.8; unadjusted p = 0.50; adjusted p = 0.21). Domain scores showed small or negligible effects throughout. ChatGPT use patterns (frequency, duration, perceived helpfulness) were not associated with performance. Participants in the intervention arm found ChatGPT helpful overall (85%), particularly for differential diagnosis (92%) and management planning (81%), but performance gains were inconsistent; 37% reported distraction. Patients expressed high acceptance and satisfaction with student ChatGPT use. Prior academic performance significantly predicted assessment scores (p = 0.04), with no preferential benefit for weaker students. Group performance in a post-study summative clinical test was similar.
Conclusions:
Minimally trained, self‑directed point‑of‑care ChatGPT use did not improve bedside performance. Any benefit is likely to depend on structured training, consistent prompts or scaffolds, and clearer workflow integration. LLM integration can support, not substitute, foundational clinical competence.
