Comparing Clusterings Using Bertin's Idea
A Pilhofer1, A Gribov, A Unwin
1University of Augsburg. alexander.pilhoefer@math.uni-augsburg.de
Abstract:
Classifying a set of objects into clusters can be done in numerous ways, producing different results. They can be visually compared using contingency tables, mosaicplots, fluctuation diagrams, tableplots, (modified) parallel coordinates plots, Parallel Sets plots or circos diagrams. Unfortunately the interpretability of all these graphical displays decreases rapidly with the numbers of categories and clusterings. In his famous book A Semiology of Graphics Bertin writes "the discovery of an ordered concept appears as the ultimate point in logical simplification since it permits reducing to a single instant the assimilation of series which previously required many instants of study". Or in more everyday language, if you use good orderings you can see results immediately that with other orderings might take a lot of effort. This is also related to the idea of effect ordering, that data should be organised to reflect the effect you want to observe. This paper presents an efficient algorithm based on Bertin's idea and concepts related to Kendall's t, which finds informative joint orders for two or more nominal classification variables. We also show how these orderings improve the various displays and how groups of corresponding categories can be detected using a top-down partitioning algorithm. Different clusterings based on data on the environmental performance of cars sold in Germany are used for illustration. All presented methods are available in the R package extracat which is used to compute the optimized orderings for the example dataset.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
The Representativeness Heuristic
Causes of Similarity-Dissimilarity Effect
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Behrens–Fisher Test
This test...
Trait Centrality


