Related Experiment Video
Updated: Jun 11, 2026

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
On Consistency and Sparsity for Principal Components Analysis in High Dimensions
Iain M Johnstone1, Arthur Yu Lu
1Iain M. Johnstone is Professor of Statistics and Biostatistics, Stanford University, Department of Statistics, 390 Serra Mall, Stanford, CA 94305 ( imj@stanford.edu ). Arthur Yu Lu is Principal, Renaissance Technologies LLC, 600 Route 25A, East Setauket, NY 11733.
Principal Component Analysis (PCA) struggles with high-dimensional data where variables exceed observations. This study proposes an initial dimensionality reduction using sparse representations to improve PCA's consistency.
Area of Science:
- Statistics
- Machine Learning
- Data Science
Background:
- Principal Component Analysis (PCA) is a standard technique for dimensionality reduction.
- Modern datasets frequently exhibit more variables (p) than observations (n), posing challenges for traditional PCA.
- Standard PCA's consistency in estimating principal components degrades when p is comparable to or exceeds n.
Purpose of the Study:
- To investigate the limitations of standard PCA in high-dimensional settings (p >> n).
- To propose and validate a preprocessing step for enhancing PCA's performance on such datasets.
- To ensure the consistency of principal component estimation even when dimensionality is very high.
Main Methods:
- Development of a simple asymptotic model to analyze PCA consistency.
- Introduction of an initial dimensionality reduction strategy using sparse representations.
- An algorithm for selecting a subset of coordinates based on largest sample variances.
Main Results:
- Standard PCA's leading component estimation is consistent if and only if p(n)/n approaches 0.
- Applying PCA to a subset of coordinates with largest variances recovers consistency.
- This subset selection method enables consistent PCA even when p(n) is much larger than n (p(n) ⪢ n).
Conclusions:
- Initial dimensionality reduction is crucial for high-dimensional data before applying PCA.
- Employing a sparse representation basis for initial reduction is an effective strategy.
- The proposed subset selection algorithm enhances PCA's reliability in p >> n scenarios.
Related Concept Videos
Principal Moments of Area
The principal moment of inertia axes are the...
Vector Algebra: Method of Components
In many applications, the magnitudes and directions of...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Friedman Two-way Analysis of Variance by Ranks

