Related Experiment Video
Updated: Jun 9, 2025

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
ChatGPT's quality: Reliability and validity of concept inventory items.
Stefan Küchemann1, Martina Rau2, Albrecht Schmidt3
1Chair of Physics Education Research, Faculty of Physics, Ludwig-Maximilians-Universität München (LMU Munich), Munich, Germany.
Large language models (LLMs) can generate physics concept items, but quality requires careful prompt engineering and expert review. Human oversight is crucial for effective educational assessments using AI-generated content.
Area of Science:
- Physics Education Research
- Artificial Intelligence in Education
Background:
- Large language models (LLMs) present opportunities and challenges in education.
- Concerns exist regarding the quality and student overreliance on LLM-generated content.
- This study evaluates LLM-generated conceptual items in physics education.
Purpose of the Study:
- To assess the quality and characteristics of conceptual physics items generated by ChatGPT.
- To compare AI-generated items with established concept inventories.
- To understand the implications for educators using LLMs in assessment creation.
Main Methods:
- Optimized prompts to generate 30 conceptual items in kinematics using ChatGPT.
- Expert review and selection of the top 15 items.
- Administered items alongside the Force Concept Inventory (FCI) to 172 university students.
- Performed confirmatory factor analysis on student responses.
Main Results:
- ChatGPT-generated items demonstrated medium difficulty and discrimination.
- AI-generated items showed slightly lower average performance compared to the FCI.
- Confirmatory factor analysis supported a three-factor model aligned with expert expectations.
Conclusions:
- High-quality conceptual items can be generated by LLMs with significant prompt engineering and selection efforts.
- AI-generated items approached the quality of human-created items but require careful vetting.
- Human oversight and student feedback are essential for refining AI-generated assessments, especially for distractors.
More Related Videos
Related Concept Videos
Reliability and Validity
Self-Report Tests of Personality
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Goodness-of-Fit Test
Quality Assurance

