Related Experiment Video
Updated: May 9, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Generalized data thinning using sufficient statistics
Ameer Dharamshi1, Anna Neufeld2, Keshav Motwani1
1Department of Biostatistics, University of Washington.
This study introduces a generalized data thinning strategy to decompose random variables into independent ones. This method expands applicability and unifies thinning with sample splitting through sufficiency.
Area of Science:
- Statistics
- Probability Theory
- Statistical Inference
Background:
- Traditional methods for decomposing random variables can fail in certain inference tasks.
- Prior work demonstrated data thinning for specific natural exponential families, requiring a summation constraint.
Purpose of the Study:
- To develop a general strategy for decomposing a random variable into independent random variables.
- To relax the summation requirement of previous thinning methods.
- To unify data thinning and sample splitting under the principle of sufficiency.
Main Methods:
- Generalizing the procedure of thinning random variables.
- Relaxing the summation constraint to a functional reconstruction.
- Applying generalized thinning to diverse statistical families.
Main Results:
- Expanded the range of distributions amenable to thinning.
- Demonstrated that data thinning and sample splitting are unified applications of sufficiency.
- Developed a general strategy applicable to a wider array of statistical families.
Conclusions:
- The generalized thinning procedure offers a more flexible approach to random variable decomposition.
- Sufficiency is identified as the unifying principle behind data thinning and sample splitting.
- The method enhances capabilities for model validation and inference.
More Related Videos
03:35Determining Gender-Based Differences in Retinal and Choroidal Thickness in Underweight Individuals via Swept-Source Optical Coherence Tomography
Published on: December 1, 2023
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Related Concept Videos
Central Limit Theorem
The sample size, n, that...
One-Way ANOVA: Unequal Sample Sizes
Quantifying and Rejecting Outliers: The Grubbs Test
Choosing Between z and t Distribution
Trimmed Mean
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...