Related Experiment Video
Updated: May 26, 2026

A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
Published on: September 4, 2019
A laboratory study on the reliability estimations of the mini-CEX
Alberto Alves de Lima1, Diego Conde, Juan Costabel
1Instituto Cardiovascular de Buenos Aires, Blanco Encalada 1525, 1428 Ciudad de Buenos Aires, Buenos Aires, Argentina. aealvesdelima@fibertel.com.ar
Reliability of workplace assessments like the mini-Clinical Evaluation Exercise (mini-CEX) is often unclear. This study used a controlled design to find that unexplained error and assessor differences significantly impact mini-CEX reliability, suggesting rater training is key.
Area of Science:
- Medical Education
- Assessment and Evaluation
- Psychometrics
Background:
- Workplace-based assessments, such as the mini-Clinical Evaluation Exercise (mini-CEX), are crucial for evaluating resident performance.
- Traditional reliability estimations often rely on real-life data, which can be confounded by factors like patient variability and assessor bias.
- The assumption of local independence in measurement is difficult to achieve in practice, leading to potential inaccuracies in reliability estimates.
Purpose of the Study:
- To estimate the reproducibility of the mini-CEX using a controlled experimental setup.
- To disentangle the sources of variance contributing to measurement error in the mini-CEX.
- To provide evidence-based recommendations for optimizing mini-CEX reliability.
Main Methods:
- A fully crossed, two-facet generalizability design was employed with 21 residents and 3 assessors.
- Each resident's three encounters were videotaped with identical patients across all residents.
- All assessors evaluated all encounters for all residents, allowing for a systematic analysis of variance.
Main Results:
- The general error term accounted for the largest proportion of variance (34%), followed by the main effect of assessors (18%).
- Universe score variance accounted for 28% of the total variance.
- Generalizability coefficients suggested that 9 encounters are needed for reliable estimation in typical practice, with fewer encounters required when using multiple assessors.
Conclusions:
- Unexplained general error and assessor leniency/stringency are primary contributors to mini-CEX unreliability.
- The findings highlight the need for strategies to mitigate these sources of error.
- Rater training is proposed as a potential intervention to enhance the reliability of the mini-CEX.
Related Concept Videos
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Contaminants and Errors
Another key consideration is determining the appropriate number of samples required to...

