Related Experiment Video
Updated: Jun 17, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Evaluating Binary Outcome Classifiers Estimated from Survey Data
Adway S Wadekar1, Jerome P Reiter
1From the Department of Statistical Science, Duke University, Durham, NC.
Using survey weights improves predictive model evaluation on complex survey data. Weighted metrics accurately reflect population performance, unlike unweighted metrics, especially with class imbalance mitigation.
Area of Science:
- Epidemiology
- Health Sciences
- Social and Behavioral Sciences
Background:
- Surveys are vital research tools but often use complex sampling designs, not simple random samples.
- Survey respondents are typically assigned weights to account for unequal selection probabilities.
- Evaluating predictive models on survey data requires careful consideration of these complex designs.
Purpose of the Study:
- To demonstrate the benefit of using survey weights for assessing predictive model quality.
- To compare weighted versus unweighted performance metrics on complex survey data.
- To evaluate the impact of weighting on models trained with class imbalance mitigation.
Main Methods:
- Characterized model assessment statistics (e.g., sensitivity, specificity) as finite population quantities.
- Computed survey-weighted estimates using random subsets of original survey data for testing.
- Conducted simulations using data from the National Survey on Drug Use and Health and National Comorbidity Survey.
Main Results:
- Unweighted metrics using sample test data can inaccurately represent population performance.
- Weighted metrics appropriately adjust for complex sampling designs, providing accurate population estimates.
- The benefit of weighted metrics persists even when models are trained using upsampling for class imbalance.
Conclusions:
- Survey weights are crucial for accurate predictive model performance evaluation on complex survey data.
- Weighted metrics provide a more reliable assessment of model generalizability to the target population.
- Researchers should adopt weighted metrics when evaluating models trained or tested on complex survey datasets.
Related Concept Videos
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Receiver Operating Characteristic Plot
Expected Frequencies in Goodness-of-Fit Tests
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Comparing the Survival Analysis of Two or More Groups
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...

