Related Experiment Videos
Scale-Invariant, Robust Sparse PCA for Large Data via Differentiable Penalties
Manu Aggarwal1, Vipul Periwal1
1Laboratory of Biological Modeling, NIDDK, National Institutes of Health, Bethesda, MD, USA.
Abstract:
Sparse PCA finds low-dimensional structure that loads on few features. Existing methods couple learning and feature selection by applying non-differentiable penalties that force retain-or-zero decisions during optimization. This eliminates features before the optimizer has established which ones matter, limits scalability to serial coordinate-update solvers, and fails when components share support. We introduce DROSS-PCA (Differentiable RObust Scalable Sparse PCA), which decouples learning and feature selection. It learns by optimizing a fully differentiable objective combining robust reconstruction, a smooth sparsity penalty, and an orthogonality term. It selects features post-hoc via cosine-preserving pruning. DROSS-PCA is GPU-accelerated, scaling sparse PCA to large data sets beyond the reach of existing methods. All loss terms are calibrated to be independent of data dimensionality and scale, giving penalty weights consistent meaning across data sets. On synthetic benchmarks, DROSS-PCA matches or outperforms established methods on support recovery. Under outlier contamination it substantially exceeds non-robust methods and remains competitive with dedicated robust methods, which do not scale to large problems. The smooth formulation of DROSS-PCA enables stability analysis via random initializations since no feature is eliminated during training. We find that the method's feature selection is stable when the generating basis is theoretically identifiable and shows increased uncertainty precisely when it is not. On the Human Lung Cell Atlas (584,944 × 27,402), gene selection is 97% consistent across independent random initializations, and consensus gene modules identified from reproducibility across runs are enriched for known biological pathways. On daily sea-surface temperature fields (ERA5, 23,376 × 484,778), the leading sparse mode reproduces the observed El Niño index and the leading modes localize to individual ocean basins.
Related Concept Videos
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear.
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Application of Linearization and Approximation
Calibration Curves: Linear Least Squares
For data that follow a straight line, the standard method for fitting is the linear...
Linear Approximations
Linearization and Approximation