Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Stratified Sampling Method01:16

Stratified Sampling Method

12.4K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
12.4K
Prediction Intervals01:03

Prediction Intervals

2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y. 
2.3K
Survival Tree01:19

Survival Tree

132
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
132
Sampling Plans01:23

Sampling Plans

241
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
241
Cluster Sampling Method01:20

Cluster Sampling Method

12.3K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.3K
Expected Frequencies in Goodness-of-Fit Tests01:19

Expected Frequencies in Goodness-of-Fit Tests

2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n)  to the number of categories (k).
2.6K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Have pulmonary function testing rates recovered post-COVID-19 pandemic? A population-based study.

Respiratory medicine·2026
Same author

Prescription patterns of inhaler medications from 2017 to 2023: A retrospective study using Ontario administrative healthcare data.

PloS one·2026
Same author

Trends in pulmonary exercise testing utilization after the COVID-19 pandemic in Ontario: A population-cohort study.

PloS one·2026
Same author

Nonparametric estimation of the total treatment effect with multiple outcomes in the presence of terminal events.

Biometrics·2026
Same author

Comparing regular expression and machine learning approaches to predict immigrant status from primary care electronic medical record data in Ontario, Canada.

PLOS digital health·2026
Same author

Data resource profile: a nationally representative linked pregnancy cohort in Canada integrating clinical, social, and environmental data.

International journal of population data science·2026

Related Experiment Video

Updated: Aug 24, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K

Efficient Evaluation of Prediction Rules in Semi-Supervised Settings under Stratified Sampling.

Jessica Gronsbell1,2, Molei Liu1,2, Lu Tian1

  • 1Jessica Gronsbell is an Assistant Professor in the Department of Statistical Sciences, University of Toronto, Toronto, ON M5S 3G3, CA Molei Liu is a Ph.D. student in the Department of Biostatistics, Harvard University, Boston, MA 02115, USA Lu Tian is an Associate Professor, Department of Biomedical Data Science, Stanford University, Palo Alto, California 94305, U.S.A Tianxi Cai is a Professor, Department of Biostatistics, Harvard University, Boston, MA 02115, USA.

Journal of the Royal Statistical Society. Series B, Statistical Methodology
|October 24, 2022
PubMed
Summary

This study introduces a novel semi-supervised learning (SSL) method for stratified sampling, improving prediction accuracy with limited labeled data. The proposed approach enhances model evaluation and efficiency in real-world applications.

Keywords:
Model EvaluationRisk PredictionSemi-Supervised LearningStratified Sampling

More Related Videos

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.2K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.6K

Related Experiment Videos

Last Updated: Aug 24, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K
An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.2K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.6K

Area of Science:

  • Machine Learning
  • Statistical Modeling
  • Data Science

Background:

  • Limited labeled data is common in many applications, driving interest in semi-supervised learning (SSL).
  • Existing SSL methods primarily assume uniform random sampling, which is often not the case in real-world scenarios.
  • There is a lack of SSL methods for performance evaluation under stratified sampling.

Purpose of the Study:

  • To propose a novel two-step SSL procedure for evaluating prediction rules under stratified sampling.
  • To address the challenge of estimating prediction performance when labeled data is not uniformly sampled.
  • To improve efficiency and consistency in semi-supervised prediction evaluation.

Main Methods:

  • A two-step semi-supervised learning procedure is proposed.
  • Step I involves imputing missing labels using weighted regression with nonlinear basis functions to handle stratified sampling.
  • Step II augments imputations for estimator consistency, regardless of prediction or imputation model specification.

Main Results:

  • The proposed SSL methods outperform supervised counterparts in efficiency gains.
  • Asymptotic theory supports the proposed methods.
  • Numerical studies demonstrate the effectiveness of the proposed approach.

Conclusions:

  • The developed SSL procedure effectively evaluates prediction rules under stratified sampling.
  • The methods offer significant efficiency improvements over traditional supervised approaches.
  • The approach is validated through an electronic health record study on diabetic neuropathy.