在高维度中交叉验证的基于损失的共变矩阵估计器选择
Philippe Boileau1, Nima S Hejazi2, Mark J van der Laan3
1Graduate Group in Biostatistics and Center for Computational Biology, UC Berkeley.
概括
在高维度中选择最好的协差矩阵估计器是具有挑战性的. 本研究引入了一种交叉验证方法,以最佳选择估计器,证明其在模拟和现实数据分析中的有效性.
科学领域:
- 统计 统计 统计 统计
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 样本共变矩阵在低维设置中是最佳的,但对于高维数据不够.
- 对于高维情景,存在许多替代的协差矩阵估计器.
- 在现有选项中选择最佳估计器仍然是一个重大挑战.
研究的目的:
- 开发一种原则方法,用于在高维设置中选择最佳的协差矩阵估计器.
- 在这种情况下,建立基于损失的交叉验证估计的理论基础.
- 提供一种实用的解决方案,用于在多种协差矩阵估计器中进行选择.
主要方法:
- 使用交叉验证的基于损失的估计框架.
- 为协差矩阵估计提出一个损失函数的一般类.
- 建立有限样本风险极限和条件,以实现交叉验证选择器的异常最佳性.
主要成果:
- 在各种数据生成过程中的数值实验中证明了拟议的交叉验证选择器的最佳性.
- 在中等样本大小中验证了该程序的有效性.
- 在使用单细胞转录组测序数据的尺寸缩小应用中展示了实际的好处.
结论:
- 拟议的交叉验证程序为选择最佳协差矩阵估计器提供了强大且理论上健全的方法.
- 这种方法解决了高维统计的关键挑战,提高了下游分析的可靠性.
- 该方法具有显著的实用优势,特别是在复杂的生物数据应用中,如单细胞RNA测序.
相关概念视频
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
589
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
589
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Goodness-of-Fit Test
3.6K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.6K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Friedman Two-way Analysis of Variance by Ranks
256
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
256


