Related Experiment Video
Updated: Sep 11, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Constructing a binary prediction model with incomplete data: Variable selection to balance fairness and precision
He Ren1, Chun Wang1, Gongjun Xu2
1Measurement and Statistics Program, College of Education, University of Washington.
Abstract:
The statistical and pragmatic tension between explanation and prediction is well recognized in psychology. Yarkoni and Westfall (2017) suggested focusing more on predictions, which will ultimately produce better calibrated interpretations. Variable selection methods, such as regularization, are strongly recommended because it will help construct interpretable models while optimizing prediction accuracy. However, when the data contain a nonignorable proportion of missingness, variable selection and model building via penalized regression methods are not straightforward. What further complicates the analysis protocol is when the model performance is evaluated on both prediction accuracy and fairness, the latter is of increasing attention when the predictive outcome has societal implications. This study explored two methods for variable selection with incomplete data: the bootstrap imputation-stability selection (BI-SS) method and the stacked elastic net (SENET) method. Both methods work with multiply imputed data sets but in different ways. BI-SS implements variable selection separately on each imputed bootstrap data set and aggregates the results via stability selection, while SENET stacks all imputed data sets and fits a single pooled model. We thoroughly evaluated their performance using a suite of metrics (including area under the curve, F1 score, and fairness criteria) via three increasingly complex simulation studies. Results reveal that while BI-SS and SENET methods perform almost equally well in settings with generalized linear models, only BI-SS fares well with nested data design because of high computation demand in fitting the regularized generalized linear mixed effects models. Finally, we demonstrated both methods with an example using rich electronic health data. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Survival Tree
Building a Survival Tree
Constructing a...
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Toward the Mean

