Related Experiment Video
Updated: Jul 4, 2025

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
Sparse Reduced Rank Huber Regression in High Dimensions
Kean Ming Tan1, Qiang Sun2, Daniela Witten3
1Department of Statistics, University of Michigan, Ann Arbor, MI.
We introduce a novel sparse reduced rank Huber regression method for high-dimensional data analysis with heavy-tailed noise. This approach offers improved statistical bias analysis and error bounds, outperforming existing methods.
Area of Science:
- Statistics
- Machine Learning
- Data Science
Background:
- High-dimensional data analysis presents challenges due to noise and complexity.
- Existing reduced rank regression methods often overlook heavy-tailed noise characteristics.
- Robust statistical methods are crucial for reliable analysis of complex datasets.
Purpose of the Study:
- To develop a robust regression method for high-dimensional data with heavy-tailed noise.
- To establish theoretical guarantees for the proposed method's estimation accuracy.
- To analyze the trade-off between noise properties and statistical bias.
Main Methods:
- Proposing a sparse reduced rank Huber regression.
- Employing convex relaxation of a non-convex optimization problem.
- Utilizing block coordinate descent and alternating direction method of multipliers algorithms.
Main Results:
- Established non-asymptotic estimation error bounds under Frobenius and nuclear norms.
- Quantified the trade-off between noise heavy-tailedness and statistical bias.
- Demonstrated convergence rates dependent on noise moment bounds, matching sub-Gaussian rates for second-moment bounded noise.
Conclusions:
- The proposed sparse reduced rank Huber regression effectively handles high-dimensional data with heavy-tailed noise.
- Theoretical analysis provides crucial insights into the method's performance under varying noise conditions.
- Numerical studies and a data application validate the method's practical utility.
Related Concept Videos
Friedman Two-way Analysis of Variance by Ranks
Quantifying and Rejecting Outliers: The Grubbs Test
Regression Toward the Mean
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Routh-Hurwitz Criterion II
The first scenario occurs when a singular zero appears in the first column of the Routh table. This situation creates a division by zero issues. To resolve this, a small positive or negative number, denoted as epsilon (∈), is substituted for the zero. The stability analysis proceeds by assuming a sign for ∈. If ∈ is positive, any sign change in the first...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...

