对于重尾分布的最佳共变性清理:信息理论的见解
Christian Bongiorno1, Marco Berritta2
1Université Paris-Saclay, CentraleSupélec, Laboratoire de Mathématiques et Informatique pour la Complexité et les Systèmes, 91192 Gif-sur-Yvette, France.
Physical review. E
|December 20, 2023
概括
最佳协差清理理论将弗罗贝尼乌斯规范最小化与正常分布的信息损失联系起来. 对于重尾学生t分布的偏差在大矩阵中减少,扩展随机矩阵理论的应用.
科学领域:
- 统计 统计 统计 统计
- 随机矩阵理论 随机矩阵理论
- 估计理论 估计理论
背景情况:
- 最佳的协同变量清理理论涉及将真和估计的协同变量矩阵之间的弗罗贝尼乌斯规范最小化.
- 旋转不变估计器在没有先前知识的情况下,对大型协同变量矩阵以非对称的方式得到.
- 学生的t分布,在金融和物理中很常见,表现出重的尾巴.
研究的目的:
- 为了证明Frobenius规范最小化和正常多变量变量信息损失之间的等价性.
- 在有限尺寸矩阵中研究学生的t分布对这个等价值的偏差.
- 探索这些偏差的非对称行为及其对随机矩阵理论的影响.
主要方法:
- 使用弗罗贝尼乌斯规范和信息丢失指标进行共变性清理的理论分析.
- 对于正常与Student的t分布的估计属性的比较.
- 对于大维共变矩阵的非对称分析.
主要成果:
- 对于正常分布,Frobenius规范最小化和信息丢失之间的等价性.
- 对有限大小的Student's t分布观察到的偏差,其中最小的弗罗贝尼乌斯规范不能保证最小的信息损失.
- 这些偏差异性地消失,这表明随机矩阵理论对Student的t分布的适用性.
结论:
- 这项研究将统计随机矩阵理论与物理学中的信息理论估计联系起来.
- 研究结果表明,随机矩阵理论的结果可能延伸到像Student的t这样的重尾分布.
- 这项工作将最佳共变性清理理论和信息理论应用联系起来.
相关概念视频
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Quantifying and Rejecting Outliers: The Grubbs Test
1.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Probability Histograms
11.6K
A probability histogram is a visual representation of a probability distribution. Similar a typical histogram, the probability histogram consists of contiguous (adjoining) boxes. It has both a horizontal axis and a vertical axis. The horizontal axis is labeled with what the data represents. The vertical axis is labeled with probability. Each rectangular bar in the histogram is 1 unit wide, which suggests that the area under each bar equals the probability, P(x), where x is 1, 2, 3, and so on.
11.6K
Chebyshev's Theorem to Interpret Standard Deviation
4.2K
Chebyshev’s theorem, also known as Chebyshev’s Inequality, states that the proportion of values of a dataset for K standard deviation is calculated using the equation:
4.2K
Calibration Curves: Correlation Coefficient
1.6K
In a linear calibration curve, there is a value called the calibration coefficient, denoted by 'r,' which measures the strength and the direction of association between two variables. The correlation coefficient value ranges from −1 to +1. A value of +1 indicates a perfect positive linear correlation, −1 denotes a perfect negative correlation, and 0 implies no correlation between the two variables. A positive correlation value establishes that as one variable increases, the...
1.6K
Probability Distributions
7.1K
The probability of a random variable x is the likelihood of its occurrence. A probability distribution represents the probabilities of a random variable using a formula, graph, or table. There are two types of probability distribution– discrete probability distribution and continuous probability distribution.
A discrete probability distribution is a probability distribution of discrete random variables. It can be categorized into binomial probability distribution and Poisson...
A discrete probability distribution is a probability distribution of discrete random variables. It can be categorized into binomial probability distribution and Poisson...
7.1K


