On Consistency and Sparsity for Principal Components Analysis in High Dimensions

Iain M Johnstone1, Arthur Yu Lu

  • 1Iain M. Johnstone is Professor of Statistics and Biostatistics, Stanford University, Department of Statistics, 390 Serra Mall, Stanford, CA 94305 ( imj@stanford.edu ). Arthur Yu Lu is Principal, Renaissance Technologies LLC, 600 Route 25A, East Setauket, NY 11733.

Summary

Principal Component Analysis (PCA) struggles with high-dimensional data where variables exceed observations. This study proposes an initial dimensionality reduction using sparse representations to improve PCA's consistency.

Related Concept Videos

Principal Moments of Area01:14

Principal Moments of Area

In mechanics, the product of inertia and moments of inertia of area help to calculate the stability and performance of various structures and components. The coordinate transformation relations are used to calculate the moments and products of inertia for an area about the inclined axes. Further, the moments and products of inertia with respect to the principal axes can be determined using the moments and products of inertia about the inclined axes.
The principal moment of inertia axes are the...
Vector Algebra: Method of Components01:08

Vector Algebra: Method of Components

It is cumbersome to find the magnitudes of vectors using the parallelogram rule or using the graphical method to perform mathematical operations like addition, subtraction, and multiplication. There are two ways to circumvent this algebraic complexity. One way is to draw the vectors to scale, as in navigation, and read approximate vector lengths and angles (directions) from the graphs. The other way is to use the method of components.
In many applications, the magnitudes and directions of...
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
One-Way ANOVA: Equal Sample Sizes01:15

One-Way ANOVA: Equal Sample Sizes

One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Variability: Analysis01:11

Variability: Analysis

Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
Friedman Two-way Analysis of Variance by Ranks01:21

Friedman Two-way Analysis of Variance by Ranks

Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures from...