Related Experiment Video
Updated: Jul 5, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Global considerations in hierarchical clustering reveal meaningful patterns in data
Roy Varshavsky1, David Horn, Michal Linial
1School of Computer Science and Engineering, The Hebrew University of Jerusalem, Jerusalem, Israel. roy.varshavsky@mail.huji.ac.il
Global hierarchical clustering methods, such as top-down (TD) and glocal algorithms, reveal more meaningful data patterns than traditional bottom-up (BU) approaches. These global methods, including a novel density-based TD algorithm, improve data analysis across various domains.
Area of Science:
- Data science
- Bioinformatics
- Machine learning
Background:
- Hierarchical clustering organizes data using tree-like structures.
- Bottom-up (BU) agglomerative algorithms are commonly used by default in unsupervised machine learning.
- Existing methods may not fully capture complex data relationships.
Purpose of the Study:
- To evaluate the effectiveness of global hierarchical clustering algorithms (top-down and glocal) compared to traditional bottom-up methods.
- To introduce and assess a novel density-based top-down hierarchical clustering algorithm.
- To demonstrate the utility of hierarchical clustering for improving data annotation and identifying errors.
Main Methods:
- Comparison of top-down (TD), glocal, and bottom-up (BU) hierarchical clustering algorithms.
- Evaluation of algorithm performance using statistical criteria based on expert-labeled data.
- Application of algorithms to diverse datasets, including gene expression, stock data, and protein families.
- Development and testing of a novel TD algorithm based on genuine data point density.
Main Results:
- Top-down (TD) and glocal algorithms are better suited for revealing meaningful patterns than BU algorithms.
- TD algorithms excel at identifying global patterns, while BU algorithms are advantageous for finer data granularity.
- The novel density-based TD algorithm outperformed existing divisive and agglomerative methods.
- Application to protein sequences identified overlooked functional annotations.
Conclusions:
- Global approaches like TD and glocal algorithms should be considered in exploratory clustering.
- Unsupervised clustering methods can enhance manually created data mappings, such as protein families.
- Hierarchical clustering provides insights into erroneous and missed annotations, improving data quality.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Outliers and Influential Points
In- and Out-Groups
Ordinal Level of Measurement
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks in the...
