Related Experiment Video
Updated: Sep 19, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessment of Large Language Model Performance on Medical School Essay-Style Concept Appraisal Questions: Exploratory
Seysha Mehta1, Eliot N Haddad1, Indira Bhavsar Burke2
1Cleveland Clinic Lerner College of Medicine, School of Medicine, Case Western Reserve University, 9500 Euclid Ave, G10, Cleveland, OH, 44195, United States, 1 2164456512, 1 2164451007.
Microsoft Copilot, a ChatGPT 4.0 large language model, performed similarly to medical students on essay questions. Assessors found it difficult to distinguish AI-generated text from human responses, highlighting the need for AI preparedness in education.
Area of Science:
- Medical Education
- Artificial Intelligence
Background:
- The integration of artificial intelligence (AI) into academic settings presents new challenges and opportunities.
- Large language models (LLMs) like ChatGPT 4.0 are increasingly capable of generating human-like text.
Purpose of the Study:
- To evaluate the performance of Microsoft Copilot (based on ChatGPT 4.0) against medical students in essay-style concept appraisals.
- To assess the ability of human evaluators to differentiate AI-generated responses from those written by students.
Main Methods:
- An essay-style concept appraisal task was administered.
- Responses were evaluated by human assessors, who were unaware of the origin of the text (AI vs. human).
Main Results:
- Microsoft Copilot demonstrated performance comparable to that of medical students.
- Assessors experienced significant difficulty in distinguishing between AI-generated and human-written responses.
Conclusions:
- AI tools like Microsoft Copilot show advanced capabilities in academic tasks, rivaling human performance.
- Educational institutions must adapt to the rise of AI by emphasizing critical thinking and reflective learning for both students and educators.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
Related Concept Videos
Language and Cognition
Patient-centered Care