相互信息和应急表的编码
Maximilian Jerdee1, Alec Kirkley2,3,4, M E J Newman1,5
1University of Michigan, Ann Arbor, Department of Physics, Michigan 48109, USA.
Physical review. E
|February 7, 2025
概括
本研究引入了一种改进的方法来计算减少的相互信息,这是数据标签之间的相似性的衡量标准. 新方法通过更好地计算应急表信息成本来纠正传统方法中的偏差,从而产生更准确的结果.
科学领域:
- 信息理论是信息理论.
- 机器学习是机器学习.
- 数据分析数据分析
背景情况:
- 相互信息是比较对象标签在分类和社区检测中的标准度量.
- 传统的相互信息计算可能会因为忽视了应急表的信息成本而产生偏见.
- 减少相互信息旨在纠正这种偏见,但依赖于准确估计信息成本限制.
研究的目的:
- 解决对标签进行比较的相互信息计算中的偏差.
- 开发一种改进的编码应急表的方法,以更好地限制信息成本.
- 为了提高减少相互信息的准确性,作为相似度衡量.
主要方法:
- 开发了一种用于应急表的新编码方法.
- 实施并测试了与传统方法相比改进的编码方法.
- 进行了广泛的数值模拟来评估性能.
主要成果:
- 改进的编码方法在典型的场景中可以更好地限制信息成本.
- 当标签非常相似时,增强的减少相互信息接近理想值.
- 数字结果证明了拟议方法的优越性.
结论:
- 新的应急表编码方法显著提高了减少相互信息的准确性.
- 这一进步为竞争性标签提供了更可靠的相似度衡量标准.
- 这些发现对于依赖于准确分类和社区检测性能量化的应用至关重要.
相关概念视频
Contingency Table
2.4K
A contingency table provides a way of portraying data that can facilitate calculating probabilities. It is a method of displaying a frequency distribution as a table with rows and columns to show how two variables may be dependent (contingent) upon each other; The table helps determine conditional probabilities quite quickly and can help systematically organize, analyze and quantify data. The table displays sample values concerning two variables that may be dependent or contingent on one...
2.4K
Introduction to Test of Independence
2.2K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
2.2K
Determination of Expected Frequency
2.1K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.1K
Friedman Two-way Analysis of Variance by Ranks
137
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
137
Hypothesis Test for Test of Independence
3.5K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
3.5K
McNemar's Test
131
McNemar's Test is a nonparametric statistical test used to determine if there is a significant difference in proportions between two related groups when the outcome is binary (e.g., yes/no, success/failure). It is beneficial when we have paired data, such as pre-test/post-test designs, where the same subjects are measured under two different conditions. The test is named after the statistician Quinn McNemar, who introduced it in 1947. It is commonly used in situations where subjects are...
131


