Related Experiment Video
Updated: Feb 17, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A novel case-control subsampling approach for rapid model exploration of large clustered binary data.
Stephen T Wright1,2,3, Louise M Ryan1,2, Tung Pham2,4
1Mathematical and Physical Sciences, University of Technology Sydney, Australia.
This study introduces a novel subsampling method for analyzing clustered binary data, enabling faster model exploration on standard workstations. The technique approximates full cohort analysis results for risk factor identification, such as in blood donation adverse reactions.
Area of Science:
- Biostatistics
- Data Science
- Epidemiology
Background:
- Identifying factors associated with outcomes is crucial for inference and prediction.
- Big data analysis, especially with complex models like Generalized Linear Mixed Models (GLMMs), can be computationally intensive and time-consuming on standard hardware.
- Efficient model exploration is needed to overcome computational constraints in big data settings.
Purpose of the Study:
- To propose a novel subsampling scheme for rapid model exploration of clustered binary data.
- To enable the use of flexible and complex model setups, such as GLMMs with additive smoothing splines, on typical workstations.
- To approximate model estimates from full cohort analyses using a case-control-type design and sampling fractions.
Main Methods:
- A novel subsampling scheme is proposed for clustered binary data.
- Prospective cohort studies are reframed into a case-control-type design.
- Cluster-specific sampling fractions are derived to incorporate cluster variation.
- The method allows for the use of GLMMs with additive smoothing splines for flexible modeling.
Main Results:
- The proposed subsampling scheme enables rapid model exploration on standard workstations.
- Model estimates from the subsampling approach approximate those from full cohort analyses.
- Computationally prohibitive analyses can be conducted in a timely manner.
- The approach was applied to analyze risk factors for adverse reactions in blood donation.
Conclusions:
- The novel subsampling method significantly speeds up the analysis of clustered binary data.
- Complex statistical models can be efficiently explored and applied on standard computing resources.
- This approach facilitates timely identification of risk factors in large datasets, exemplified by blood donation adverse reactions.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Statistical Methods for Analyzing Epidemiological Data
Quantifying and Rejecting Outliers: The Grubbs Test
Comparing the Survival Analysis of Two or More Groups
Survival Tree
Building a Survival Tree
Constructing a...

