Related Experiment Video
Updated: Jun 12, 2026

Validation of a Psychosocial Intervention on Body Image in Older People: An Experimental Design
Published on: May 31, 2021
Secure AI-assisted angoff standard-setting for single best answer questions: A non-inferiority validation study
Edward Stephenson1,2, Sarah Robinson1, Kate Bascombe1
1Department of Medical Education, Brighton & Sussex Medical School, Brighton, UK.
AI-powered Angoff estimates for medical exams are comparable to human judgments, offering a faster and less burdensome alternative for setting pass scores. While AI shows promise, human experts remain essential for final decisions.
Area of Science:
- Medical Education
- Artificial Intelligence in Assessment
Background:
- The Angoff method for setting exam pass scores is expert-driven, time-consuming, and prone to variability.
- Evaluating AI-driven Angoff estimates offers a potential solution to improve efficiency and consistency.
Purpose of the Study:
- To assess if AI-derived Angoff estimates are non-inferior to human judgments for single best answer (SBA) questions.
- To explore the utility of a secure, offline AI tool for feature extraction in standard-setting.
Main Methods:
- A methodological study in a UK Physician Associate program utilized a borderline student descriptor and an AI tool (ExamFeats) to extract features from SBAs.
- Three AI models (LLM, ML, Hybrid) were compared against average human Angoff ratings for 100 new SBAs, with a 10% non-inferiority margin.
Main Results:
- AI models produced mean Angoff estimates closely aligned with human scores (60.0%-60.8% vs. 60.3% human).
- No significant differences were found between AI models and human ratings, with confidence intervals within the non-inferiority margin.
- Discordance occurred in 33% of SBAs, not explained by item features, suggesting AI should augment, not replace, human expertise.
Conclusions:
- AI-derived Angoff estimates can replicate panel-level behavior within acceptable error bounds, potentially reducing the burden of standard setting.
- A secure feature-extraction tool allows AI-assisted standard setting without compromising item security.
- Question-level discordance highlights the need for AI to supplement, rather than substitute, expert judgment in exam standard setting.
Related Concept Videos
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Self-Evaluation: Self-Enhancement and Self-Verification
One-Way ANOVA: Unequal Sample Sizes
Strategies of Self-Presentation II: Self-Verification
Bonferroni Test
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...