Related Experiment Video
Updated: Apr 7, 2026

Using Learning Outcome Measures to assess Doctoral Nursing Education
Published on: June 21, 2010
Can Large Language Models Replicate Systematic Review Outcome Classifications in Medical Education? A Pilot Study
Giuliano Romano1, Emilio Romano2, Michelle Rau1
1Oakland University William Beaumont School of Medicine, Rochester, MI United States of America.
ChatGPT showed modest agreement when classifying medical education outcomes using the Kirkpatrick framework. Further research is needed to improve AI
Area of Science:
- Medical Education
- Artificial Intelligence in Healthcare
- Systematic Review Methodology
Background:
- Systematic reviews in medical education commonly use the Kirkpatrick framework for outcome classification.
- Manual coding of outcomes is labor-intensive and prone to subjectivity.
- Advancements in AI offer potential solutions for automating such tasks.
Purpose of the Study:
- To evaluate the performance of ChatGPT (GPT-5, August 2025 release) in classifying outcomes within medical education systematic reviews.
- To assess the agreement between AI-driven classification and human expert coding using the Kirkpatrick framework.
Main Methods:
- A proof-of-concept study was performed using 32 full-text articles from a systematic review on sepsis education.
- ChatGPT was employed to code outcomes based on the Kirkpatrick framework.
- Agreement was quantified using percentage agreement, unweighted kappa, and weighted kappa statistics.
Main Results:
- The agreement between ChatGPT and human coders was modest, with 50% percent agreement.
- Unweighted kappa (κ) was 0.170 (95% CI 0.000-0.458) and weighted kappa (κ) was 0.351 (95% CI 0.074-0.629).
- Most discrepancies occurred between adjacent levels of the Kirkpatrick framework.
Conclusions:
- ChatGPT demonstrates potential for assisting in the classification of outcomes in medical education systematic reviews.
- Current performance indicates limitations, with a need for refinement to improve accuracy and consistency.
- Further development and validation are necessary before widespread adoption of AI tools for systematic review data extraction.
More Related Videos
13:44Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques
Published on: December 9, 2022
10:39Qualitative and Quantitative Validation of Tools with Rating Scales Aimed at Assessing the Quality of University Service-Learning
Published on: August 29, 2025
Related Concept Videos
Nursing Process for Patient and Caregiver Teaching III: Evaluation and Documentation
Nurses can use several methods to evaluate patient outcomes. For example, oral questions can assess cognitive learning,...
Types of Records II: Educational and Administrative Records
Nursing Evaluation
Guidelines for Writing Outcome
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care...
Levels of Organization
Molecules Are Composed of Atoms, and Biomolecules Are Assembled from Molecules:
The most basic levels include atoms, molecules, and biomolecules. Atoms, the smallest unit of ordinary matter, are composed of a nucleus and electrons. Molecules...
Heart Failure IV: Classification and Diagnostic Evaluation