Benchmarking ChatGPT-generated multiple-choice questions against faculty-authored items in dental education
Amber Kiyani1, Fariha Hanif2, Muhammad Muhammad2
1Qatar University, Doha, Qatar. a.kiyani@qu.edu.qa.
This study compared ChatGPT-generated multiple-choice questions (MCQs) with instructor-written ones for dental students. Instructor items performed better, but AI-generated questions show future potential in educational assessments.
Area of Science:
- Educational Technology
- Artificial Intelligence in Education
- Medical Education
Background:
- Large language models (LLMs) are increasingly used in self-directed student learning.
- Assessing the quality of AI-generated educational content is crucial for effective integration.
- Benchmarking AI tools against traditional methods provides insights into their utility.
Purpose of the Study:
- To compare the psychometric properties of assessment items generated by ChatGPT-3.5 against those created by faculty instructors.
- To evaluate the performance of AI-generated multiple-choice questions (MCQs) in an Oral Medicine context for undergraduate dental students.
- To analyze the reliability, difficulty, and discrimination of both AI-generated and instructor-created assessment items.
Main Methods:
- A 40-item Oral Medicine assessment was developed, comprising 20 MCQs by instructors and 20 by ChatGPT-3.5.
- The assessment was administered to 547 undergraduate dental students across four institutions.
- Item Response Theory, specifically the 3-Parameter Logistic model, was used for data analysis.
Main Results:
- Instructor items demonstrated higher person reliability (0.845) compared to ChatGPT items (0.778).
- Instructor items showed a wider difficulty range (-1.29 to 3.22) and higher discrimination (0.50 to 4.64) than ChatGPT items.
- While instructor items had higher average scores and variability, ChatGPT items exhibited promising psychometric properties with a low guessing parameter.
Conclusions:
- Faculty-generated assessment items currently exhibit superior psychometric properties, including better discrimination and difficulty range.
- ChatGPT-generated items show potential for MCQ generation in educational settings, warranting further investigation.
- LLMs like ChatGPT may play a significant role in future educational assessment development, potentially surpassing human capabilities.
More Related Videos
13:44Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques
Published on: December 9, 2022
07:42Quasistatic Mechanical Testing for Computer-Aided Design and Manufacturing Occlusal Veneers Cemented to Milled Dentin Analog Material
Published on: December 20, 2024
Related Concept Videos
Assessment of the Mouth
Mouth Inspection
The inspection begins with visually examining the mouth for symmetry, color, and size.
Surveys
Biostatistics: Overview
Discrete variables are...
The Availability Heuristic
Comparing Experimental Results: Student's t-Test
Pre-Procedural Guidelines for Assessing Blood Pressure
