Related Experiment Video
Updated: May 29, 2026

Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size (LEfSe) in Microbiome Data
Published on: May 16, 2022
The use of shrinkage estimators in linear discriminant analysis
1Programs in Mathematical Sciences, University of Texas at Dallas, Richardson, TX 75080.
Abstract:
Probably the most common single discriminant algorithm in use today is the linear algorithm. Unfortunately, this algorithm has been shown to frequently behave poorly in high dimensions relative to other algorithms, even on suitable Gaussian data. This is because the algorithm uses sample estimates of the means and covariance matrix which are of poor quality in high dimensions. It seems reasonable that if these unbiased estimates were replaced by estimates which are more stable in high dimensions, then the resultant modified linear algorithm should be an improvement. This paper studies using a shrinkage estimate for the covariance matrix in the linear algorithm. We chose the linear algorithm, not because we particularly advocate its use, but because its simple structure allows one to more easily ascertain the effects of the use of shrinkage estimates. A simulation study assuming two underlying Gaussian populations with common covariance matrix found the shrinkage algorithm to significantly outperform the standard linear algorithm in most cases. Several different means, covariance matrices, and shrinkage rules were studied. A nonparametric algorithm, which previously had been shown to usually outperform the linear algorithm in high dimensions, was included in the simulation study for comparison.
Related Concept Videos
Shrinkage in Concrete
When concrete is still in its plastic state, it can undergo a decrease in volume by about 1% of its absolute volume. This decrease is known as plastic shrinkage. It arises either...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
What are Estimates?
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such as the mean,...
Drying Shrinkage
A portion of this drying shrinkage can be reversed; if the concrete is...
Calibration Curves: Linear Least Squares
For data that follow a straight line, the standard method for fitting is the linear...
Linearization and Approximation
