Related Experiment Video
Updated: Aug 26, 2026

Inverse Probability of Treatment Weighting (Propensity Score) using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Machine learning models for chronic disease risks using complex survey data: the impact of sample weights
Hyeonju Kim1, Paul Rogers1, Dong Wang1
1Division of Bioinformatics and Biostatistics, National Center for Toxicological Research, U.S. Food and Drug Administration, Jefferson, AR, United States.
None:
Chronic noncommunicable diseases are the leading cause of deaths globally but can be preventable by early detection, screening, and treatment. Recent studies have used machine learning methods ignoring sample weights determined by a complex design framework, which leads to significant bias and ends up with the spurious conclusion in statistical inference. We have examined whether adjusting sample weights improves prediction and variable selection using two examples, hypertension and diabetes, from the National Health and Nutrition Examination Survey data. We used the traditional K-fold cross validation and its weighted counterparts, replicate weights methods, to finalize models fitted by four popular machine learning methods and two under-sampling approaches for handling class imbalance ad hoc. Due to lack of options managing complex sampling structures in existing machine learning software packages, We have also developed a r-MLSurvey that integrates the replicate weights methods into the machine learning methods as an extension of weighted L1-penalized logistic regression. The model performance was assessed by accuracy, sensitivity, specificity, the area under the receiver operating characteristic curve (AUC-ROC), and their counterparts for the weighted model evaluation. Results show that variable selection results differ by the under-sampling methods and sample weights' adjustment; the weighted AUC, compared with the AUC, suggests final unweighted models either overestimate or underestimate the true population parameters. The sampling design also appears to provide moderate influence on accuracy for these relatively well-studied diseases, but the effect on selected variables' significance varies by algorithm and class imbalance. Random forest would be preferred without users' capability of adjusting sample weights. Careful selection of the final models is recommended for moderately class-imbalanced complex survey data to avoid biasedness. Our application to the machine learning algorithms offers the usefulness of complex survey data to achieve unbiased prediction in broader research disciplines in need of sufficiently large data.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Statistical Methods for Analyzing Epidemiological Data