基于量子PSO,在社交网络中使用属性对用户进行无监督的集群
Debadatta Naik1, Ramesh Dharavath1, Lianyong Qi2
1Indian Institute of Technology (ISM), Dhanbad, India.
概括
本研究介绍了一种新的量子PSO集群方法,用于社交网络,重点关注用户属性. 它通过提高聚类准确性和克服局部最佳值来改善用户组发现,从而改进了K模式.
科学领域:
- 社交网络分析 社交网络分析
- 数据挖掘 数据挖掘
- 机器学习 机器学习
背景情况:
- 无监督的集群检测在社交网络中将类似的用户组合在一起.
- 现有的方法经常使用链接或链接和属性.
- 基于属性的聚类是有价值的,但K模式可以面对局部最佳.
研究的目的:
- 提出一种新的聚类方法,只使用用户属性.
- 解决K模式算法在社交网络分析中的局限性.
- 为了提高社交网络用户聚类的准确性.
主要方法:
- 基于属性的缩小维度 (属性的选择和删除).
- 量子粒子群集优化 (QPSO) 用于最大化相似性.
- 使用了三种不同的相似度,用于属性减少和聚类.
主要成果:
- 拟议的量子PSO方法展示了优越的集群性能.
- 在ego-Twitter和ego-Facebook数据集上表现优于K-Mode和K-Mean算法.
- 在三个关键绩效指标上取得了更好的结果.
结论:
- 量子PSO方法有效地检测基于用户属性的社交网络集群.
- 这种方法为分类数据的传统集群算法提供了改进的替代方案.
- 该方法提升了用户相似性最大化,以实现更准确的社交网络细分.
相关概念视频
Cluster Sampling Method
12.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.0K
Outliers and Influential Points
4.1K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.1K
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
99
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
99
Quantifying and Rejecting Outliers: The Grubbs Test
1.7K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.7K
Pore Size Distribution
170
In concrete, the pore size distribution significantly influences the material's properties. Capillary pores, markedly larger than gel pores, form a vast network within partially hydrated cement paste, reducing the concrete's strength and increasing its permeability. This heightened permeability leads to a greater risk of damage from environmental factors like freeze-thaw cycles and chemical attacks, with the extent of vulnerability also being tied to the water-to-cement ratio.
Adequate...
Adequate...
170
Classification of Systems-I
221
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
221


