Related Experiment Video
Updated: May 22, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Optimal variable clustering for high-dimensional matrix valued data
Inbeom Lee1, Siyi Deng2, Yang Ning3
1Booth School of Business, University of Chicago, 5807 S. Woodlawn Ave., Chicago, IL 60637, USA.
This study introduces a novel latent variable model and hierarchical clustering algorithm for matrix-valued data, leveraging feature dependence structures. The proposed method achieves high-dimensional clustering consistency and optimal performance, outperforming existing techniques.
Area of Science:
- Statistics
- Machine Learning
- Data Science
Background:
- Matrix-valued data is increasingly common in various applications.
- Existing clustering methods often focus on the mean model and neglect the informative feature dependence structure.
- This limitation is particularly relevant in high-dimensional settings or when mean information is insufficient.
Purpose of the Study:
- To develop a new latent variable model for matrix-valued data that utilizes the feature dependence structure for clustering.
- To propose hierarchical clustering algorithms based on a weighted covariance matrix dissimilarity measure.
- To theoretically analyze the clustering consistency and optimality of the proposed method.
Main Methods:
- A novel latent variable model for matrix-valued data is proposed, incorporating row and column membership matrices.
- A class of hierarchical clustering algorithms is developed using the difference of a weighted covariance matrix as a dissimilarity measure.
- Theoretical analysis includes establishing clustering consistency in high-dimensional settings and deriving minimax lower bounds.
Main Results:
- The proposed algorithm demonstrates clustering consistency under mild conditions in high-dimensional settings.
- An optimal weight for the covariance matrix is identified, ensuring minimax rate-optimality.
- Simulation studies show superior performance compared to existing methods, evidenced by a higher adjusted Rand index (ARI).
Conclusions:
- The developed latent variable model and hierarchical clustering algorithm effectively utilize the dependence structure of matrix-valued data.
- The method provides theoretical guarantees for high-dimensional clustering and achieves optimal performance.
- The approach offers practical advantages and yields meaningful interpretations, as demonstrated on a genomic dataset.
Related Concept Videos
Friedman Two-way Analysis of Variance by Ranks
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Outliers and Influential Points
Column Efficiency: Rate Theory
During elution, a solute molecule experiences numerous transitions between stationary and mobile phases, exhibiting irregular residence times in...
Coefficient of Correlation
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...

