规范化的相互信息是对分类和社区检测的一种有偏见的措施
Maximilian Jerdee1,2, Alec Kirkley3,4,5, Mark Newman6,7
1Department of Physics, University of Michigan, Ann Arbor, MI, USA.
Nature communications
|December 11, 2025
概括
在集群和分类评估中,规范化的相互信息 (NMI) 是有偏见的. 本研究引入了修改后的NMI来纠正这些偏差,显著影响了算法性能的结论,特别是网络社区检测.
科学领域:
- 计算机科学 计算机科学
- 信息理论 信息理论
- 数据挖掘 数据挖掘
背景情况:
- 规范化相互信息 (NMI) 是用于评估聚类和分类算法的标准度量.
- 现有的NMI计算包含了由于忽视信息内容和对称规范化的偏差而产生的偏差.
研究的目的:
- 解决传统NMI的局限性.
- 引入一个修改后的,公正的相互信息措施.
- 为了证明NMI偏差对算法性能结论的影响.
主要方法:
- 开发了一个修改后的相互信息计算.
- 对网络社区检测算法进行了广泛的数值测试.
- 使用传统的NMI与修改措施的结果进行比较.
主要成果:
- 在传统的NMI中发现了两个关键偏差:忽视应急表信息和对对称规范化的虚假依赖.
- 修改后的相互信息化措施纠正了这些发现的缺陷.
- 关于网络社区检测表现最好的算法的结论通过使用不偏见的测量来显著改变.
结论:
- 传统的NMI在绩效评估中引入了显著的偏见.
- 拟议的修改后的相互信息提供了一个更准确,更可靠的相似度.
- 在算法比较研究中,特别是在网络社区检测中,使用公正的测量对于得出正确的结论至关重要.
相关概念视频
Mutual Inductance
3.5K
Inductance is the property of a device that tells us how effectively it induces an emf in another device. In other words, it is a physical quantity that expresses the effectiveness of a given device.
When two circuits carrying time-varying currents are close to one another, the magnetic flux through each circuit varies because of the changing current in the other circuit. Consequently, an emf is induced in each circuit by the changing current in the other. Therefore, this type of emf is called...
When two circuits carrying time-varying currents are close to one another, the magnetic flux through each circuit varies because of the changing current in the other circuit. Consequently, an emf is induced in each circuit by the changing current in the other. Therefore, this type of emf is called...
3.5K
Confidence Coefficient
10.3K
The confidence coefficient is also known as the confidence level or degree of confidence. It is the percent expression for the probability, 1-α, that the confidence interval contains the true population parameter assuming that the confidence interval is obtained after sufficient unbiased sampling; for example, if the CL = 90%, then in 90 out of 100 samples the interval estimate will enclose the true population parameter. Here α is the area under the curve, distributed equally under...
10.3K
Midrange
4.1K
A somewhat easy to compute quantitative estimate of a data set’s central tendency is its midrange, which is defined as the mean of the minimum and maximum values of an ordered data set.
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
4.1K
Test for Homogeneity
2.3K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.3K
Kendall's Coefficient of Concordance
915
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
915
Aggregates Classification
950
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
950


