Related Experiment Video
Updated: May 9, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Generalized data thinning using sufficient statistics
Ameer Dharamshi1, Anna Neufeld2, Keshav Motwani1
1Department of Biostatistics, University of Washington.
Abstract:
Our goal is to develop a general strategy to decompose a random variable into multiple independent random variables, without sacrificing any information about unknown parameters. A recent paper showed that for some well-known natural exponential families, can be thinned into independent random variables , such that . These independent random variables can then be used for various model validation and inference tasks, including in contexts where traditional sample splitting fails. In this paper, we generalize their procedure by relaxing this summation requirement and simply asking that some known function of the independent random variables exactly reconstruct . This generalization of the procedure serves two purposes. First, it greatly expands the families of distributions for which thinning can be performed. Second, it unifies sample splitting and data thinning, which on the surface seem to be very different, as applications of the same principle. This shared principle is sufficiency. We use this insight to perform generalized thinning operations for a diverse set of families.
More Related Videos
03:35Determining Gender-Based Differences in Retinal and Choroidal Thickness in Underweight Individuals via Swept-Source Optical Coherence Tomography
Published on: December 1, 2023
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Related Concept Videos
Central Limit Theorem
The sample size, n, that...
One-Way ANOVA: Unequal Sample Sizes
Quantifying and Rejecting Outliers: The Grubbs Test
Choosing Between z and t Distribution
Trimmed Mean
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...