Related Experiment Video
Updated: Apr 27, 2026

05:21
Computerized Adaptive Testing System of Functional Assessment of Stroke
Published on: January 7, 2019
5.4K
A comparison of stopping rules for computerized adaptive screening measures using the rating scale model
Audrey J Leroux1, Barbara G Dodd
1University of Texas at Austin, 1 University Station D4900, Austin, TX 78712, USA, audreyleroux@utexas.edu.
Summary
The predicted standard error reduction (PSER) stopping rule is best for identifying at-risk students with computerized adaptive testing (CAT). This method improves classification accuracy and reduces testing time, especially with smaller item pools.
Area of Science:
- Educational Measurement
- Psychometrics
- Computerized Adaptive Testing (CAT)
Background:
- Computerized adaptive testing (CAT) uses adaptive algorithms to tailor tests to individual ability levels.
- Selecting appropriate stopping rules is crucial for balancing test efficiency and measurement precision in CAT.
- Identifying at-risk students requires accurate and efficient assessment methods.
Purpose of the Study:
- To evaluate and compare three stopping rules for CAT: predicted standard error reduction (PSER), fixed-length, and minimum SE.
- To assess the effectiveness of these rules in classifying at-risk students using Andrich's rating scale model.
- To examine the impact of variables like trait distribution and item pool size on classification accuracy.
Main Methods:
- Simulated data were generated using Andrich's rating scale model.
- Three stopping rules (PSER, fixed-length, minimum SE) were applied to the simulated CAT data.
- Classification accuracy for at-risk students was evaluated under various conditions (trait distribution, item pool size).
Main Results:
- The PSER stopping rule demonstrated superior performance in correctly classifying at-risk students.
- PSER was effective in reducing the number of items administered, thereby alleviating test burden.
- The benefits of PSER were particularly evident with smaller item pools and specific trait distributions.
Conclusions:
- The PSER stopping rule is recommended for CAT applications focused on identifying at-risk students.
- PSER offers a balance between accurate classification and efficient testing, minimizing student burden.
- This finding is especially relevant for screening measures utilizing rating scale models with limited item sets.
Related Concept Videos
Response Surface Methodology
889
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
The process of RSM involves several key steps:
889
Friedman Two-way Analysis of Variance by Ranks
594
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
594
Receiver Operating Characteristic Plot
568
A ROC (Receiver Operating Characteristic) plot is a graphical tool used to assess the performance of a binary classification model by illustrating the trade-off between sensitivity (true positive rate) and specificity (false positive rate). By plotting sensitivity against 1 - specificity across various threshold settings, the ROC curve shows how well the model distinguishes between classes, with a curve closer to the top-left corner indicating a more accurate model. The area under the ROC curve...
568
Measures of Intelligence
12.7K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
12.7K

