Compositional data analysis tutorial
Michael Smithson1, Stephen B Broomell2
1Research School of Psychology, Australian National University.
Abstract:
This article presents techniques for dealing with a form of dependency in data arising when numerical data sum to a constant for individual cases, that is, "compositional" or "ipsative" data. Examples are percentages that sum to 100, and hours in a day that sum to 24. Ipsative scales fell out of fashion in psychology during the 1960s and 1970s due to a lack of methods for analyzing them. However, ipsative scales have merits, and compositional data commonly occur in psychological research. Moreover, as we demonstrate, sometimes converting data to a compositional form yields insights not otherwise accessible. Fortunately, there are sound methods for analyzing compositional data. We seek to enable researchers to analyze compositional data by presenting appropriate techniques and illustrating their application to real data. First, we elaborate the technical details of compositional data and discuss both established and new approaches to their analysis. We then present applications of these methods to real social science data-sets (data and code using R are available in a supplementary document). We conclude with a discussion of the state of the art in compositional data analysis and remaining unsolved problems. A brief guide to available software resources is provided in the first section of the supplementary document. (PsycInfo Database Record (c) 2024 APA, all rights reserved).
Related Concept Videos
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Classifying Matter by Composition
According to its composition, the matter can be classified into two broad categories — pure substances and mixtures.
A pure substance is a form of matter that has a constant composition throughout with uniform properties. For example, any sample of sucrose has the same composition and same physical properties, such as melting point, color, and sweetness, regardless of the source from which it is isolated.
A mixture is composed of two or...
Two-Way ANOVA
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the...
One-Way ANOVA
Statistical Methods to Analyze Parametric Data: ANOVA
One-way ANOVA is applied when a single independent variable or factor is scrutinized. It compares...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:


