Related Experiment Video
Updated: Dec 11, 2025

12:18
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
7.8K
Bias in Estimation of Misclassification Rates.
1Educational Testing Service, Princeton. shaberman@ets.org.
Psychometrika
|February 16, 2017
Summary
Using random sampling for classification rules increases misclassification rates, but this effect is minimal in most statistical analyses. The impact lessens significantly as sample size grows, especially with polytomous variables.
Area of Science:
- Statistics
- Machine Learning
- Predictive Modeling
Background:
- Classification rules are essential for predicting outcomes based on independent variables.
- The accuracy of these rules can be affected by the sampling method used.
Purpose of the Study:
- To quantify the impact of simple random sampling on misclassification rates in polytomous variable prediction.
- To compare the misclassification rates when using estimated versus known conditional probability distributions.
Main Methods:
- Analysis of misclassification rates in predictive classification models.
- Mathematical assessment of excess error due to sampling in various predictor-response variable scenarios (polytomous-polytomous, continuous-polytomous).
Main Results:
- Simple random sampling increases the best achievable misclassification rate compared to using known distributions.
- This increase is generally small, diminishing rapidly (exponentially) with sample size for polytomous predictors.
- For continuous predictors, excess error typically scales with sample size to the power of -2/3.
Conclusions:
- While sampling introduces a slight increase in misclassification error, it is often negligible in practical statistical analysis.
- The sensitivity of misclassification error to prediction quality differs from probability-based criteria, which may introduce bias.

