通过批量自身价值匹配分析估计协差矩阵中峰值自身价值的数量
Zheng Tracy Ke1, Yucong Ma1, Xihong Lin2
1Department of Statistics of Harvard University.
Journal of the American Statistical Association
|February 27, 2025
概括
这项研究引入了一种用于估计高维数据中尖端固有值数量的新方法. 通过利用大量的固有值,新方法在统计分析中提供了更好的准确性和稳定性.
科学领域:
- 统计 统计 统计 统计
- 高维数据分析 高维数据分析
- 协方差建模的模型
背景情况:
- 尖的协差模型越来越多地用于分析高维数据.
- 估计尖峰自值 (k) 的数量至关重要,但具有挑战性.
- 现有的方法主要集中在顶部自值上,忽视了批量自值.
研究的目的:
- 开发一种使用大量固有值估计k的原则方法.
- 为了提高在尖端协差模型中自值估计的准确性和稳定性.
- 为高维数据分析提供可靠的统计工具.
主要方法:
- 在剩余共变矩阵上强加一个工作模型,假设马分布的对角线条目.
- 大量固有值用于估计固定参数分布的参数.
- 建议采用两步估计程序,利用大量的自身价值信息.
主要成果:
- 拟议的估计器 (k_hat) 汇总了来自众多批量固有值的信息.
- 估计器的一致性是在标准尖峰协差模型下证明的.
- 开发了k的置信区间,提高了估计可靠性.
结论:
- 新方法有效地结合了批量固有值,以便对k进行可靠的估计.
- 模拟研究证实了该方法对现有方法的优越性.
- 该方法通过在肺癌和基因组数据集中的应用得到了验证.
相关概念视频
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
331
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
331
Wilcoxon Signed-Ranks Test for Matched Pairs
78
The Wilcoxon signed-rank test for matched pairs evaluates the null hypothesis by combining the ranks of differences with their signs. It essentially tests whether the median of the differences in a population of matched pairs is zero. Since the test incorporates more information than the sign test, it generally yields more trustable conclusions. This test also does not require the data to follow a normal distribution, but two conditions must be met for it to be applicable: (1) the data must...
78
Quantifying and Rejecting Outliers: The Grubbs Test
1.4K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.4K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Kendall's Coefficient of Concordance
214
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
214
Calculating and Interpreting the Linear Correlation Coefficient
5.9K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
5.9K


