Related Experiment Videos
Early evidence for a multi-agent AI simulator for clinical reasoning practice: performance, consistency, and
Matthew E Kelleher1, Christine Y Zhou2, Danielle E Weber1
1Department of Pediatrics and Internal Medicine, Cincinnati Children's Hospital Medical Center and University of Cincinnati College of Medicine, Cincinnati, OH, USA.
Objectives:
Clinical reasoning develops through repeated, deliberate practice, yet clinical and simulation environments are often limited by continuity, feedback, and scalability constraints. Large language models (LLMs) may address this by generating virtual clinical encounters accessible without live instructors or standardized patients but validity evidence remains limited. This study explored the performance of MAESSCR (Multi-Agent Educational Scenario Simulator for Clinical Reasoning), a platform where multiple LLM-based agents, each assigned a distinct role (i.e. patient, physical exam, diagnostic testing) interact with learners through a text-based encounter.
Methods:
In Fall of 2024, 175 second-year medical students completed three MAESSCR clinical encounters as coursework. Six clinician-educators developed a rating tool to evaluate 120 transcripts (40 randomly sampled per case) using a dichotomous (yes/no) scale across four domains: (1) realism of agent responses, (2) adherence to scripted case details, (3) platform functionality, and (4) interference with students' independent reasoning through clinical findings. Qualitative narrative review supplemented binary ratings.
Results:
MAESSCR followed scripted details in 92 % (110/120) of transcripts. Diagnosis-changing information occurred in 2.5 % (3/120). Unrealistic patient portrayal appeared in 12 % (14/120) of encounters, and technical issues in 17 % (20/120). AI agents most often interfered with student's clinical reasoning by interpreting findings before the students had the opportunity. This occurred in 28 % (34/120) of encounters due to history/physical exam agents and 59 % (71/120) because of diagnostics/management agents. Most disruptions were minor and unlikely to compromise the overall encounter.
Conclusions:
Multi-agent, LLM-based simulations offer a scalable approach to deliberate practice of clinical reasoning. Educational value depends on role stability, contextual fidelity, and learners' opportunity to interpret clinical information independently. Specific design considerations are essential to ensure AI-generated simulations support, rather than disrupt, the clinical reasoning process.
Related Concept Videos
Critical Thinking II
Reason and Intuition
Reasoning
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Patient-centered Care