Related Experiment Video
Updated: Oct 2, 2025

Employing the Forced Oscillation Technique for the Assessment of Respiratory Mechanics in Adults
Published on: February 9, 2022
Pass/fail decisions and standards: the impact of differential examiner stringency on OSCE outcomes
1School of Medicine, Leeds Institute of Medical Education, University of Leeds, LS29JT, Leeds, UK. m.s.homer@leeds.ac.uk.
Abstract:
Variation in examiner stringency is a recognised problem in many standardised summative assessments of performance such as the OSCE. The stated strength of the OSCE is that such error might largely balance out over the exam as a whole. This study uses linear mixed models to estimate the impact of different factors (examiner, station, candidate and exam) on station-level total domain score and, separately, on a single global grade. The exam data is from 442 separate administrations of an 18 station OSCE for international medical graduates who want to work in the National Health Service in the UK. We find that variation due to examiner is approximately twice as large for domain scores as it is for grades (16% vs. 8%), with smaller residual variance in the former (67% vs. 76%). Combined estimates of exam-level (relative) reliability across all data are 0.75 and 0.69 for domains scores and grades respectively. The correlation between two separate estimates of stringency for individual examiners (one for grades and one for domain scores) is relatively high (r=0.76) implying that examiners are generally quite consistent in their stringency between these two assessments of performance. Cluster analysis indicates that examiners fall into two broad groups characterised as hawks or doves on both measures. At the exam level, correcting for examiner stringency produces systematically lower cut-scores under borderline regression standard setting than using the raw marks. In turn, such a correction would produce higher pass rates-although meaningful direct comparisons are challenging to make. As in other studies, this work shows that OSCEs and other standardised performance assessments are subject to substantial variation in examiner stringency, and require sufficient domain sampling to ensure quality of pass/fail decision-making is at least adequate. More, perhaps qualitative, work is needed to understand better how examiners might score similarly (or differently) between the awarding of station-level domain scores and global grades. The issue of the potential systematic bias of borderline regression evidenced for the first time here, with sources of error producing cut-scores higher than they should be, also needs more investigation.
More Related Videos
06:29Comparing Objective Conjunctival Hyperemia Grading and the Ocular Surface Disease Index Score in Dry Eye Syndrome During COVID-19
Published on: May 25, 2022
08:33A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
Published on: September 4, 2019
Related Concept Videos
Reliability and Validity
Detection of Gross Error: The Q Test
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Assessment of Airway, Skin Color, and Use of Accessory Muscles
Introduction
The initial evaluation of a patient's respiratory system...
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...