在非统一的缺失边缘模式下,在现实世界网络上的链接预测准确性
Xie He1, Amir Ghasemian2, Eun Lee3
1Department of Mathematics, Dartmouth College, Hanover, NH, United States of America.
PloS one
|July 18, 2024
概括
链接预测的准确性因网络数据的收集方式而异. 这项研究指导研究人员选择适合非统一的缺失数据模式的算法,这是现实世界网络中常见的.
科学领域:
- 网络科学 网络科学
- 数据挖掘 数据挖掘
- 机器学习 机器学习
背景情况:
- 现实世界网络数据集通常由于数据收集偏差而缺少边缘.
- 统一的缺失数据是评估链接预测算法的一个常见的,但往往不现实的假设.
研究的目的:
- 为了调查不同的非统一的缺失边缘模式如何影响链接预测的准确性.
- 在各种缺失数据场景下比较各种链接预测算法的性能.
- 为根据网络数据特征选择合适的链接预测方法提供指导.
主要方法:
- 分析了来自4个家族的9个链接预测算法.
- 在20个不同的缺失边缘模式中进行评估,分为5个组.
- 一项比较模拟研究,使用来自6个领域的250个现实世界网络数据集.
主要成果:
- 在不同的缺失边缘模式中观察到链接预测算法性能的显著变化.
- 该研究强调了非统一的缺失数据对评估结果的重大影响.
- 算法的性能高度依赖于缺失数据的特定特征.
结论:
- 假设统一的缺失数据可能会导致对链接预测方法的误导性评估.
- 研究人员在选择算法时应考虑数据收集过程和由此产生的缺失边缘模式.
- 这项工作为选择针对真实世界网络数据量身定制的链接预测工具提供了一个框架.
更多相关视频
10:44Inherent Dynamics Visualizer, an Interactive Application for Evaluating and Visualizing Outputs from a Gene Regulatory Network Inference Pipeline
Published on: December 7, 2021
2.1K
03:31Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
515
相关概念视频
End Point Prediction: Gran Plot
310
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
310
Survival Tree
78
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
78
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Improving Translational Accuracy
9.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.9K
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
