Assessment of Large Language Model Performance on Medical School Essay-Style Concept Appraisal Questions: Exploratory

Seysha Mehta1, Eliot N Haddad1, Indira Bhavsar Burke2

  • 1Cleveland Clinic Lerner College of Medicine, School of Medicine, Case Western Reserve University, 9500 Euclid Ave, G10, Cleveland, OH, 44195, United States, 1 2164456512, 1 2164451007.

PubMed
Summary

Microsoft Copilot, a ChatGPT 4.0 large language model, performed similarly to medical students on essay questions. Assessors found it difficult to distinguish AI-generated text from human responses, highlighting the need for AI preparedness in education.

Related Concept Videos