Related Experiment Video
Updated: Mar 26, 2026

12:27
Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
7.4K
A Graph Theoretic Criterion for Determining the Number of Clusters in a Data Set
Multivariate Behavioral Research
|January 27, 2016
Summary
This study introduces a novel criterion for determining the number of clusters in data. Based on graph theory, it offers a mathematically sound approach to solving the cluster analysis problem.
Area of Science:
- Data Science
- Computer Science
- Statistics
Background:
- Current methods for determining the number of clusters often lack formal grounding in clustering models.
- Existing stopping rules focus on cluster coherence and isolation but are not model-derived.
Purpose of the Study:
- To introduce a new criterion for identifying the optimal number of clusters in a dataset.
- To provide a mathematically rigorous approach to cluster analysis.
- To evaluate the performance of the new criterion against established methods.
Main Methods:
- Utilizing graph theoretic concepts, including minimal spanning trees and maximal spanning trees.
- Developing a novel criterion based on homomorphic functions for clustering.
- Comparing the proposed criterion with four existing stopping rules on empirical datasets.
Main Results:
- The new criterion demonstrates mathematically attractive properties.
- The proposed method effectively determines the number of clusters in various datasets.
- Performance evaluation shows the criterion's potential to solve the number-of-clusters problem.
Conclusions:
- The novel graph-theory-based criterion offers a robust solution for determining the number of clusters.
- This approach provides a mathematically sound alternative to existing, less formal stopping rules.
- The criterion has practical implications for data analysis and cluster validation.
Related Concept Videos
Quantifying and Rejecting Outliers: The Grubbs Test
4.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
4.5K
Histogram
18.8K
The histogram is a graphical representation in the x-y form of data distribution in a data set. The horizontal x-axis is labeled with what the data represents (for instance, distance from your home to school). The vertical y-axis is labeled either frequency or relative frequency (or percent frequency or probability).
A histogram graph consists of contiguous (adjoining) boxes. The heights of the bars correspond to frequency values. The graph will have the same shape with respective labels. The...
A histogram graph consists of contiguous (adjoining) boxes. The heights of the bars correspond to frequency values. The graph will have the same shape with respective labels. The...
18.8K
Bar Graph
23.7K
A bar graph is also called a bar chart and consists of bars that are separated from each other. It either uses horizontal or vertical bars to show comparisons among categories. The bars can be rectangles, or they can be rectangular boxes (used in three-dimensional plots). One axis of the graph represents the specific categories being compared, and the other axis shows a discrete value. In this graph, the length of the bar for each category is proportional to the number or percent of individuals...
23.7K
Multiple Bar Graph
10.5K
As the name suggests, a multiple bar graph is the same as a bar graph but has multiple bars to depict relationships between different data values. One can include as many parameters as possible. However, each parameter must have the same unit of measurement.
Each bar or column in the multiple bar graph represents a data value. These graphs are used primarily in interrelating two or more sets of data. The categories of different kinds of data are listed along the horizontal or x-axis, whereas...
Each bar or column in the multiple bar graph represents a data value. These graphs are used primarily in interrelating two or more sets of data. The categories of different kinds of data are listed along the horizontal or x-axis, whereas...
10.5K
Graphical Representation of Inequalities
378
The graph of the equation where y equals x squared forms a curve known as a parabola. This curve acts as a boundary in the coordinate plane, dividing it into distinct regions based on the relative position of points.When the equality sign in the equation is replaced with an inequality—such as greater than, less than, greater than or equal to, or less than or equal to—the graphical representation changes from a single curve into a broader shaded area that signifies the set of all...
378
Ogive Graph
7.1K
An ogive graph is sometimes called a cumulative frequency polygon. It is one type of frequency polygon that shows cumulative frequency. In other words, the cumulative percentages are added to the graph from left to right. An ogive graph plots cumulative frequency on the vertical y-axis and class boundaries along the horizontal x-axis. It’s very similar to a histogram; only instead of rectangles, an ogive displays a single point where the top right of the rectangle would be. Creating this...
7.1K

