稀疏图的半监督集群:越过信息理论值
Junda Sheng1, Thomas Strohmer2
1Department of Mathematics, University of California, Davis, CA 95616-5270, USA.
概括
在使用随机区块模型的网络中,社区检测在稀疏的图表中是有限的. 然而,在半监督的环境中加入即使是很小一部分标签也可以消除这一局限性,使所有参数都能准确检测.
科学领域:
- 网络科学 网络科学
- 统计推断的统计推断.
- 机器学习 机器学习
背景情况:
- 随机区块模型是网络社区检测的一个基本工具.
- 在Kesten-Stigum值存在一个关键限制,阻碍了稀疏图表的性能.
- 现有的方法在低于这个值的性能下扎.
研究的目的:
- 调查半监督学习对随机区块模型限制的影响.
- 以部分标签信息来证明社区检测的可行性.
- 开发用于整合网络结构和标签的新算法.
主要方法:
- 在半监督环境中对随机区块模型进行理论分析.
- 开发一个用于标签集成的组合算法.
- 为标签集成开发基于优化的算法.
主要成果:
- 通过部分标签,克服了凯斯坦-斯蒂格姆门所施加的根本限制.
- 在整个参数域中,社区检测变得可行.
- 引入了两个高效的算法,利用图形拓和标签数据.
结论:
- 半监督学习显著提高了随机区块模型的功能.
- 拟议的算法为在现实世界网络中的社区检测提供了实际的解决方案.
- 这项研究为网络分析和半确定的编程开辟了新的途径.
相关概念视频
Cluster Sampling Method
11.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.6K
Quantifying and Rejecting Outliers: The Grubbs Test
1.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.5K
Routh-Hurwitz Criterion II
193
In the application of the Routh-Hurwitz criterion, two specific scenarios can arise that complicate stability analysis.
The first scenario occurs when a singular zero appears in the first column of the Routh table. This situation creates a division by zero issues. To resolve this, a small positive or negative number, denoted as epsilon (∈), is substituted for the zero. The stability analysis proceeds by assuming a sign for ∈. If ∈ is positive, any sign change in the first...
The first scenario occurs when a singular zero appears in the first column of the Routh table. This situation creates a division by zero issues. To resolve this, a small positive or negative number, denoted as epsilon (∈), is substituted for the zero. The stability analysis proceeds by assuming a sign for ∈. If ∈ is positive, any sign change in the first...
193
Survival Tree
61
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
61
Routh-Hurwitz Criterion I
173
Consider an electrical power grid, where stability is essential to prevent blackouts. The Routh-Hurwitz criterion is a valuable tool for assessing system stability under varying load conditions or faults. By analyzing the closed-loop transfer function, the Routh-Hurwitz criterion helps determine whether the system remains stable.
To apply the Routh-Hurwitz criterion, a Routh table is constructed. The table's rows are labeled with powers of the complex frequency variable s, starting from the...
To apply the Routh-Hurwitz criterion, a Routh table is constructed. The table's rows are labeled with powers of the complex frequency variable s, starting from the...
173
Outliers and Influential Points
4.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.0K


