Related Experiment Videos
K-means clustering: a half-century synthesis
1Department of Psychological Sciences, University of Missouri-Columbia, Columbia, MO 65211, USA. steinleyd@missouri.edu
Summary
This paper reviews fifty years of K-means clustering research, covering its core concepts, variations, and statistical underpinnings. It provides a unified view of K-means and its extensions, highlighting areas for future study.
Area of Science:
- Data Science
- Machine Learning
- Statistical Analysis
Background:
- The K-means clustering algorithm is a fundamental technique in unsupervised machine learning.
- Its widespread application necessitates a comprehensive understanding of its theoretical basis and practical considerations.
Purpose of the Study:
- To synthesize fifty years of research on the K-means clustering method.
- To provide a unifying treatment of K-means and its various extensions.
- To identify future research directions for enhancing K-means performance.
Main Methods:
- Review and synthesis of existing literature on K-means clustering.
- Discussion of theoretical statistical results related to the algorithm.
- Exploration of different formulations, initialization techniques, and preprocessing schemes.
Main Results:
- A comprehensive overview of K-means clustering, including its core principles and loss function formulations.
- Detailed discussion of methods for selecting the number of clusters, initialization, variable preprocessing, and data reduction.
- Presentation of theoretical statistical results and extensions of the K-means algorithm using different metrics and modifications.
Conclusions:
- The paper offers a unifying perspective on K-means clustering and its related methods.
- Identifies key areas for future research to address performance subtleties.
- Emphasizes the enduring relevance and adaptability of K-means in data analysis.