Related Experiment Video
Updated: Jan 17, 2026

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
Published on: December 7, 2021
Fast Probabilistic Whitening Transformation for Ultra-High Dimensional Genetic Data
Gabriel E Hoffman1,2, Panos Roussos1,2, Kiran Girdhar1,2
1Center for Disease Neurogenomics, Department of Psychiatry, Department of Genetics and Genomic Sciences, Icahn School of Medicine at Mount Sinai, New York, NY, USA.
This study introduces a novel probabilistic model for data whitening, offering a faster and more accurate method for analyzing high-dimensional datasets. The new approach improves performance in genetic studies and disease research.
Area of Science:
- Statistics
- Machine Learning
- Bioinformatics
Background:
- Statistical methods often assume sample independence, which is frequently violated in real-world data due to ubiquitous correlation structures.
- Existing whitening transformations, based on linear algebra, struggle with high-dimensional data (p >> n) and have prohibitive cubic time complexity.
- The limitations of current methods hinder their application in fields like genomics where datasets are large and complex.
Purpose of the Study:
- To propose a probabilistic model for data whitening that addresses the limitations of existing linear algebra-based approaches.
- To develop a computationally efficient algorithm for whitening high-dimensional data.
- To evaluate the performance of the probabilistic whitening model on simulated and real-world genetic data.
Main Methods:
- Developed a probabilistic model for data whitening based on first principles.
- Derived a new algorithm with linear time complexity with respect to the number of features (p).
- Validated the model's out-of-sample performance using simulated datasets and real genotype data.
Main Results:
- The probabilistic whitening model demonstrated superior statistical properties and computational efficiency compared to traditional methods.
- Achieved the lowest mean square error in imputing z-statistics for a schizophrenia genome-wide association study, with speeds up to ten times faster.
- Identified tandem repeats linked to genetic regulatory signals for disease-relevant genes.
Conclusions:
- The proposed probabilistic whitening approach offers a statistically sound and computationally efficient alternative for analyzing correlated, high-dimensional data.
- This method has significant implications for genetic association studies and understanding disease-relevant genetic mechanisms.
- Open-source R packages 'decorrelate' and 'imputez' are available for implementing these analyses.
Related Concept Videos
Mutation, Gene Flow, and Genetic Drift
Genetic Drift
Genetic Variation
Genes exist in different versions called alleles,...
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...

