Related Experiment Video
Updated: Dec 7, 2025

Computerized Adaptive Testing System of Functional Assessment of Stroke
Published on: January 7, 2019
Similarity of the cut score in test sets with different item amounts using the modified Angoff, modified Ebel, and
Janghee Park1, Mi Kyoung Yim2, Na Jin Kim3
1Department of Medical Education, Soonchunhyang University College of Medicine, Asan, Korea.
Purpose:
The Korea Medical Licensing Exam (KMLE) typically contains a large number of items. The purpose of this study was to investigate whether there is a difference in the cut score between evaluating all items of the exam and evaluating only some items when conducting standard-setting.
Methods:
We divided the item sets that appeared on 3 recent KMLEs for the past 3 years into 4 subsets of each year of 25% each based on their item content categories, discrimination index, and difficulty index. The entire panel of 15 members assessed all the items (360 items, 100%) of the year 2017. In split-half set 1, each item set contained 184 (51%) items of year 2018 and each set from split-half set 2 contained 182 (51%) items of the year 2019 using the same method. We used the modified Angoff, modified Ebel, and Hofstee methods in the standard-setting process.
Results:
Less than a 1% cut score difference was observed when the same method was used to stratify item subsets containing 25%, 51%, or 100% of the entire set. When rating fewer items, higher rater reliability was observed.
Conclusion:
When the entire item set was divided into equivalent subsets, assessing the exam using a portion of the item set (90 out of 360 items) yielded similar cut scores to those derived using the entire item set. There was a higher correlation between panelists' individual assessments and the overall assessments.
Related Concept Videos
Bonferroni Test
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
Friedman Two-way Analysis of Variance by Ranks
Expected Frequencies in Goodness-of-Fit Tests
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Wilcoxon Signed-Ranks Test for Matched Pairs
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...

