识别网络集群质量指标中的偏差
Martí Renedo-Mirambell1, Argimiro Arratia1
1Soft Computing Research Group (SOCO) at Intelligent Data Science and Artificial Intelligence Research Center, Department of Computer Sciences, Universitat Politécnica de Catalunya, Barcelona, Spain.
PeerJ. Computer science
|September 14, 2023
概括
网络集群质量指标往往有利于更少,更大的集群. 研究人员开发了新的模型来测试这些指标,发现模块化和密度比对社区检测来说不那么有偏见.
科学领域:
- 网络科学 网络科学
- 数据分析数据分析
- 算法评价算法评价的方法
背景情况:
- 网络集群质量指标评估社区结构.
- 现有的指标可能会显示与内部和外部连接相关的偏差.
- 评估这些偏见对于准确的网络分析至关重要.
研究的目的:
- 调查流行网络集群质量指标中的潜在偏见.
- 开发一种强大的方法来生成具有受控社区结构和学位分配的网络.
- 引入和评估一种新的质量指标,即密度比.
主要方法:
- 使用随机和偏好的附件块模型来生成网络.
- 集成的预设社区结构,Poisson和无尺度的度分布.
- 创建多层结构以测试不同集群数量和强度的指标性能.
主要成果:
- 大多数评估的指标都显示了偏向偏好较少,较大的集群分区.
- 这种偏见甚至在内部和外部连接可比的情况下也存在.
- 与其他指标相比,密度比度指标显示偏差减少.
结论:
- 流行的网络集群指标通常表现出固有的偏见.
- 模块化和拟议的密度比指标似乎不太容易受到这些偏差的影响.
- 仔细选择质量指标对于网络中可靠的社区检测至关重要.
相关概念视频
Cluster Sampling Method
12.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.0K
Bias in Epidemiological Studies
343
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
343
Quantifying and Rejecting Outliers: The Grubbs Test
1.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Bias
4.3K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
4.3K
Confidence Coefficient
7.7K
The confidence coefficient is also known as the confidence level or degree of confidence. It is the percent expression for the probability, 1-α, that the confidence interval contains the true population parameter assuming that the confidence interval is obtained after sufficient unbiased sampling; for example, if the CL = 90%, then in 90 out of 100 samples the interval estimate will enclose the true population parameter. Here α is the area under the curve, distributed equally under...
7.7K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K


