A medoid-based deviation ratio index to determine the number of clusters in a dataset
Kariyam1,2, Abdurakhman1, Adhitya Ronnie Effendie1
1Department of Mathematics, Faculty of Mathematics and Natural Sciences, Gadjah Mada University, Indonesia.
Abstract:
Most existing methods of determining the number of groups apply to particular data types or are calculated based on the distance matrix for all object pairs. In this paper, we propose a medoid-based Deviation Ratio Index (DRI) to determine the number of clusters. The DRI is calculated based on the distance matrix for each object to final medoids. These final medoids are produced by the block-based -medoids algorithm (BlockD-KM). We choose a specific transformation and a suitable distance for certain variables before executing the BlockD-KM. We illustrated the detailed stages of DRI on secondary data in the 2022 environmental index of Asia Pacific countries, so that they are easy to reproduce. We use eight real datasets, namely Breast Cancer, Heart Disease, Iris, Wine, Soybean, Ionosphere, Vote, and Credit Approval data, to validate the DRI method. We compare the DRI method with the Calinski-Harabaz (CH) and the Silhouette index. The experimental results show that the DRI is 100% correct in predicting the number of clusters. While the CH index correctly predicts 62.5% and the Silhouette index of 75%. We also generated three kinds of artificial data to evaluate the proposed method, and 76.7% of the experiments were predicted correctly.•The medoid-based deviation ratio index aids the researcher in determining the number of clusters•The DRI method applicable to any medoids-based partitioning algorithm•This method is suitable for all data types (categorical, numerical, and mixed).
Related Concept Videos
Mean Absolute Deviation
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
Polymers: Molecular Weight Distribution
Chebyshev's Theorem to Interpret Standard Deviation
Midrange
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Range Rule of Thumb to Interpret Standard Deviation
For instance, the range rule of thumb can be used to find the tallest and the shortest student in a class, given the mean student height and standard deviation. If the mean student height is 1.6 m and the standard deviation, s is 0.05 m, the height...


