Related Experiment Video
Updated: Sep 9, 2026

Spatial Separation of Molecular Conformers and Clusters
Published on: January 9, 2014
A unified MAP-EM approach to stable Gaussian mixture clustering with priors, graphs, and split-merge adaptation for
Sumathi Subbarayan1, G Hannah Grace1
1Department of Mathematics, School of Advanced Sciences, Vellore Institute of Technology Chennai, Chennai, Tamil Nadu, India.
Introduction:
Clustering high-dimensional and noisy data remains challenging for conventional expectation-maximization (EM) methods as overlapping clusters, sparse features, and outliers can lead to covariance degeneracy and unstable parameter estimates. This research aims to improve clustering performance in high-dimensional, noisy settings by developing a robust maximum a posteriori expectation-maximization (MAP-EM) framework that integrates prior regularization, geometric structure, and outlier handling. Traditional EM-based clustering methods often struggle in the presence of overlapping clusters, high-dimensional features, and outliers, leading to degenerate covariance and unstable parameter estimates.
Methods:
The proposed MAP-EM model improves reliability by combining normal-inverse-Wishart (NIW) priors for covariance stabilization, a graph-Laplacian structure over the feature space to capture geometric relations among features, a uniform noise component to absorb outliers, and an adaptive split-merge strategy that refines cluster boundaries. These modules are coupled within a single MAP-EM procedure. The noise component modifies the E-step responsibilities by capturing atypical observations, and these updated responsibilities drive the NIW-regularized covariance and the graph-regularized mean updates in the M-step. The split-merge step is accepted only if it improves the penalized objective. The proposed MAP-EM model was evaluated on five synthetic datasets and three benchmark text corpora, namely Reuters-R8, BBC Sports, and the BBC dataset.
Results:
On Reuters-R8, the model achieved an Adjusted Rand Index of 0.347 and an accuracy of 0.569, outperforming the variational Bayesian Gaussian mixture model (GMM) (0.317). On the BBC dataset, it achieved the highest Adjusted Rand Index of 0.326 and a Normalized Mutual Information of 0.425 among the methods compared. Formal statistical testing showed that MAP-EM achieved significant positive differences in 45 out of 75 method-level comparisons, with one significant negative comparison. At the run level, MAP-EM obtained higher scores in 891 out of 1,115 valid paired comparisons, corresponding to a win rate of 79.9%. Theoretical analysis further supports the proposed framework by establishing coercivity of the penalized objective, monotone ascent of the MAP-EM iterations, finite termination of the split-merge stage under the stated acceptance criterion, boundedness of the covariance estimates under the normal-inverse-Wishart prior, local R-linear convergence of the MAP-EM iterations after model-order stabilization, eigenvalue bounds for the NIW covariance estimator, a condition-number bound for the penalized mean update, and a perturbation bound for the graph-regularized mean update.
Discussion:
The proposed MAP-EM framework provides a stable, structure-aware clustering approach for high-dimensional text data. Experimental results indicate that its advantages are most evident on datasets with noise, overlapping clusters, and well-connected feature graphs, rather than across all clustering scenarios.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an organic...
Multicompartment Models: Overview
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
