Variance-Based Cluster Selection Criteria in a K-Means Framework for One-Mode Dissimilarity Data.
J Fernando Vera1, Rodrigo Macías2
1Department of Statistics and O.R., Faculty of Sciences, University of Granada, 18071, Granada, Spain. jfvera@ugr.es.
Psychometrika
|February 15, 2017
Summary
Determining the number of clusters is key in cluster analysis. This study proposes new variance-based criteria for one-mode dissimilarity matrices, outperforming existing methods for finding the optimal number of groups.
Area of Science:
- Statistics
- Data Mining
- Machine Learning
Background:
- Determining the optimal number of clusters is a fundamental challenge in cluster analysis.
- Existing methods, like those for K-means, often rely on two-mode data sets and total point scatter decomposition.
- Adapting these criteria for one-mode dissimilarity matrices, where object coordinates are unknown, requires reformulation.
Purpose of the Study:
- To formulate novel criteria for determining the number of clusters using one-mode dissimilarity matrices.
- To adapt existing variance-based criteria for situations where only dissimilarities are available.
- To evaluate the performance of these new criteria compared to traditional methods.
Main Methods:
- Proposed decomposition of object variability based on block-shaped partitions of the dissimilarity matrix.
- Derived within-block and between-block dispersion values from the partitioned dissimilarity matrix.
- Formulated new variance-based criteria for cluster number determination.
- Conducted Monte Carlo simulations to assess criterion performance.
Main Results:
- The proposed variance-based criteria, when applied to Euclidean distances derived from one-mode data, showed greater efficiency in recovering the number of clusters.
- This improved efficiency was particularly notable for unequal-sized clusters and in low-dimensional settings.
- When applied directly to dissimilarity data, the proposed criteria consistently outperformed their original formulations.
Conclusions:
- The developed variance-based criteria offer an effective approach for determining the number of clusters from one-mode dissimilarity matrices.
- Reformulating criteria using derived Euclidean distances can enhance cluster recovery accuracy, especially in challenging data configurations.
- These findings provide valuable tools for researchers dealing with dissimilarity data in cluster analysis.
Related Concept Videos
One-Way ANOVA: Equal Sample Sizes
4.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
4.3K
Cluster Sampling Method
15.3K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
15.3K
What is Variation?
19.1K
Apart from the measures of central tendency, distribution, outliers, and the changing characteristics of data with time, an important characteristic of any data set is its variation or spread. In some data sets, the data values are concentrated closely near the mean; in others, the data values are more widely spread out from the mean.
The range, standard deviation, standard error, and variance are the different measures of variation.
Range: The range is the difference between its maximum and...
The range, standard deviation, standard error, and variance are the different measures of variation.
Range: The range is the difference between its maximum and...
19.1K
Causes of Similarity-Dissimilarity Effect
307
The similarity-dissimilarity effect, a fundamental concept in social psychology, explains how interpersonal similarities and differences influence attraction and social interactions. This effect is supported by three key psychological perspectives: balance theory, social comparison theory, and consensual validation.Balance Theory and Cognitive ConsistencyBalance theory, developed by Fritz Heider, posits that individuals seek cognitive consistency in their relationships. When two people share...
307
One-Way ANOVA
14.0K
One-way ANOVA analyzes more than three samples categorized by one factor. For example, it can compare the average mileage of sports bikes. Here, the data is categorized by one factor - the company. However, one-way ANOVA cannot be used to simultaneously compare the sample mean of three or more samples categorized by two factors. An example of two factors would be sports bikes from different companies driven in different terrains, such as a desert or snowy landscape. Here, two-way ANOVA is used...
14.0K
One-Way ANOVA: Unequal Sample Sizes
6.8K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
6.8K


