Related Experiment Video
Updated: Mar 28, 2026

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
A Comparison Study on Similarity and Dissimilarity Measures in Clustering Continuous Data
Ali Seyed Shirkhorshidi1, Saeed Aghabozorgi2, Teh Ying Wah1
1Department of Information Systems, Faculty of Computer Science and Information Technology, University of Malaya, 50603, Kuala Lumpur, Malaysia.
This study introduces a framework to evaluate similarity measures in high-dimensional data for clustering. It benchmarks common measures, guiding researchers to select appropriate methods for diverse datasets.
Area of Science:
- Data Science
- Machine Learning
- Computational Statistics
Background:
- Similarity and distance measures are crucial for distance-based clustering algorithms.
- Existing research primarily focuses on low-dimensional spaces (2D-3D), leaving a gap in understanding performance in high-dimensional datasets.
Purpose of the Study:
- To propose a technical framework for analyzing, comparing, and benchmarking the influence of various similarity measures on distance-based clustering results.
- To address the lack of empirical studies on similarity measure performance in high-dimensional spaces.
Main Methods:
- Development of a technical framework to systematically evaluate similarity measures.
- Utilizing fifteen publicly available datasets, categorized into low and high-dimensional groups.
- Benchmarking the performance of different similarity measures within these categories.
Main Results:
- The study provides an empirical analysis of similarity measure behavior in high-dimensional clustering.
- Established a benchmark for evaluating existing and novel distance measures.
- Identified how different measures perform across low- and high-dimensional datasets.
Conclusions:
- The proposed framework facilitates the selection of appropriate distance measures for specific datasets.
- Enables reproducible comparison and evaluation of newly proposed similarity or distance measures against traditional ones.
- Contributes empirical evidence on the performance of similarity measures in high-dimensional data clustering.
More Related Videos
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
07:59Author Spotlight: Alignment of Synchronized Time-Series Data Using the Characterizing Loss of Cell Cycle Synchrony Model for Cross-Experiment Comparisons
Published on: June 9, 2023
Related Concept Videos
Causes of Similarity-Dissimilarity Effect
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
What is Variation?
The range, standard deviation, standard error, and variance are the different measures of variation.
Range: The range is the difference between its maximum and...
Comparing the Survival Analysis of Two or More Groups
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Test for Homogeneity