Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Stratified Sampling Method01:16

Stratified Sampling Method

12.3K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
12.3K
Systematic Sampling Method01:17

Systematic Sampling Method

10.6K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
Systematic sampling is one of the simplest methods...
10.6K
Cluster Sampling Method01:20

Cluster Sampling Method

12.2K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.2K
Bootstrapping01:24

Bootstrapping

658
The term "bootstrap" originated in the 19th century as a metaphor for self-improvement or achieving something independently, without external assistance. This concept extends to statistical bootstrapping, a self-contained method for estimating population parameters through resampling, even though it can be computationally intensive. Developed by the American statistician Dr. Bradley Efron in 1979, bootstrapping provides a robust way to perform inference when the original sample size is...
658
Sampling Plans01:23

Sampling Plans

233
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
233
Sampling Methods: Overview01:06

Sampling Methods: Overview

405
A sample refers to a smaller subset representative of a larger population. In analytical chemistry, studying or analyzing an entire population is often impractical or impossible. Therefore, samples are used to draw inferences and generalize the whole population. The sampling method selects individuals or items from a population to create a sample. Standard sampling methods include random, judgemental, systematic, stratified, and cluster sampling. 
In analytical chemistry, the choice of...
405

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Corrigendum to "Revisiting the modifiable areal unit problem in the era of exposome-wide association studies: Assessing the performance of the CDC/ATSDR social vulnerability index at privacy-protecting spatial scales" [Environ. Res. (2026) 124912].

Environmental research·2026
Same author

Shared trans-ancestry architecture of HLA-mediated disease risk in the <i>All of Us</i> Research Program.

medRxiv : the preprint server for health sciences·2026
Same author

Prenatal Smoking Exposures and Epigenome-Wide Methylation in Newborn Blood.

Environmental health perspectives·2026
Same author

Corrigendum to "Revisiting the modifiable areal unit problem in the era of exposome-wide association studies: Assessing the performance of the CDC/ATSDR social vulnerability index at privacy-protecting spatial scales" [Environ. Res. (2026) 124912].

Environmental research·2026
Same author

COX-2-Derived PGE<sub>2</sub> Modulates IL-17 Production by γδ T Cells During Allergic Lung Inflammation.

FASEB journal : official publication of the Federation of American Societies for Experimental Biology·2026
Same author

Revisiting the modifiable areal unit problem in the era of exposome-wide association studies: Assessing the performance of the CDC/ATSDR social vulnerability index at privacy-protecting spatial scales.

Environmental research·2026

Related Experiment Video

Updated: Aug 14, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

14.6K

Implementing machine learning methods with complex survey data: Lessons learned on the impacts of accounting sampling

Nathaniel MacNell1, Lydia Feinstein1,2, Jesse Wilkerson1

  • 1Social & Scientific Systems, a DLH Holdings Company, Durham, North Carolina, United States of America.

Plos One
|January 13, 2023
PubMed
Summary

Ignoring sampling weights in gradient boosting models can reduce the generalizability of complex survey data predictions. Recalculating model performance using weighted outcomes may offer a more accurate assessment when weighted algorithms are unavailable.

More Related Videos

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.6K
Watershed Planning within a Quantitative Scenario Analysis Framework
12:44

Watershed Planning within a Quantitative Scenario Analysis Framework

Published on: July 24, 2016

8.1K

Related Experiment Videos

Last Updated: Aug 14, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

14.6K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.6K
Watershed Planning within a Quantitative Scenario Analysis Framework
12:44

Watershed Planning within a Quantitative Scenario Analysis Framework

Published on: July 24, 2016

8.1K

Area of Science:

  • Epidemiologic research
  • Machine learning
  • Survey data analysis

Background:

  • Machine learning methods are increasingly used in epidemiology, but software often lacks support for complex survey data.
  • Analyzing complex survey data requires accounting for sampling weights for valid predictions in target populations.

Purpose of the Study:

  • To determine if ignoring sampling weights in gradient boosting models affects prediction accuracy for all-cause mortality.
  • To assess the impact of sample size, weight variability, predictor strength, and model dimensionality on prediction accuracy.

Main Methods:

  • Utilized data from 15,820 participants in the National Health and Nutrition Examination Survey (1988-1994).
  • Employed gradient boosting models to predict all-cause mortality, comparing weighted and unweighted analyses.
  • Conducted simulations to evaluate various factors influencing prediction accuracy.

Main Results:

  • Unweighted models in the NHANES data showed inflated F1 scores (81.9%) compared to weighted models (77.4%).
  • Recalculating unweighted model performance using weighted outcomes mitigated this error (F1: 77.0%).
  • In simulations, this mitigation effect was consistent for large sample sizes (N=10,000) but varied for smaller samples.

Conclusions:

  • Failing to account for sampling weights in gradient boosting models can limit the generalizability of findings from complex survey data.
  • Post-hoc recalculations of model performance using weighted observed outcomes can provide a more accurate prediction assessment when weighted algorithms are unavailable.