Identifying Key Nodes for the Influence Spread Using a Machine Learning Approach
Mateusz Stolarski1, Adam Piróg2, Piotr Bródka1
1Department of Artificial Intelligence, Wrocław University of Science and Technology, 50-370 Wrocław, Poland.
Entropy (Basel, Switzerland)
|November 27, 2024
Summary
This study introduces "Smart Bins" to improve machine learning for identifying key influencers in complex networks. The enhanced framework accurately predicts influence spread and generalizes across various network types.
Area of Science:
- Network Science
- Machine Learning
- Complex Systems Analysis
Background:
- Identifying key nodes is crucial for applications like viral marketing and epidemic control.
- Machine learning (ML) methods show promise but require refinement for accuracy and generalization.
Purpose of the Study:
- To develop an enhanced ML framework for identifying key nodes in complex networks, specifically for the Independent Cascade model.
- To address challenges in obtaining training labels and improving model generalization.
Main Methods:
- Introduction of a novel "Smart Bins" technique for improved label generation in ML training.
- Development of an ML-based framework to predict influence spread and node characteristics.
Main Results:
- "Smart Bins" demonstrate superiority over existing methods for generating training labels.
- The proposed framework accurately predicts node influence and reveals additional spreading process characteristics.
- Extensive testing confirms the framework's robust generalization across diverse network structures and sizes.
Conclusions:
- The enhanced ML framework offers a significant advancement in identifying key nodes for influence spread prediction.
- The "Smart Bins" method and the framework's ability to extract further spreading characteristics represent novel contributions to network science.
Related Concept Videos
Outliers and Influential Points
4.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.0K
End Point Prediction: Gran Plot
278
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
278
Cluster Sampling Method
11.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.6K
Confidence Coefficient
7.5K
The confidence coefficient is also known as the confidence level or degree of confidence. It is the percent expression for the probability, 1-α, that the confidence interval contains the true population parameter assuming that the confidence interval is obtained after sufficient unbiased sampling; for example, if the CL = 90%, then in 90 out of 100 samples the interval estimate will enclose the true population parameter. Here α is the area under the curve, distributed equally under...
7.5K
Survival Tree
61
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
61
Classification of Signals
403
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
403


