Related Experiment Videos
Clustering ensembles: models of consensus and weak partitions.
Alexander Topchy1, Anil K Jain, William Punch
1Nielsen Media Research, 501 Brooker Creek Blvd., Oldsmar, FL 34677, USA. alexander.topchy@nielsenmedia.com
IEEE Transactions on Pattern Analysis and Machine Intelligence
|December 17, 2005
Summary
Clustering ensembles enhance unsupervised classification stability. This study introduces a probabilistic consensus model and a new consensus function, improving accuracy even with weak clustering components and incomplete data.
Area of Science:
- Machine Learning
- Data Mining
- Computational Statistics
Background:
- Clustering ensembles improve robustness and stability in unsupervised classification.
- Consensus clustering is challenging, with existing approaches from graph-based, combinatorial, or statistical perspectives.
Purpose of the Study:
- To develop a unified representation for multiple clusterings and formulate a categorical clustering problem.
- To propose a probabilistic consensus model using finite mixture of multinomial distributions.
- To define a novel consensus function based on generalized mutual information.
Main Methods:
- Formulating a categorical clustering problem with a unified representation.
- Developing a probabilistic consensus model using finite mixture of multinomial distributions.
- Employing the EM algorithm for maximum-likelihood estimation and defining a new consensus function.
Main Results:
- Demonstrated efficacy of combining weak clustering algorithms (data projections, random splits).
- Analyzed combination accuracy based on component partition parameters and number of partitions.
- Evaluated performance with incomplete information and missing cluster labels.
Conclusions:
- The proposed probabilistic model and consensus function effectively improve clustering ensemble performance.
- The methods are robust to weak clustering components and handle incomplete data scenarios.
- Experimental results validate the effectiveness on real-world datasets.