Omics data in relative values are almost subcompositionally coherent
Marina Martínez-Álvaro1, Michael Greenacre2, Agustín Blasco1
1Instituto de Ciencia y Tecnología Animal, Universitat Politècnica de València, Valencia, Spain.
Introduction:
Omics data are compositional and often expressed as relative abundances after total sum scaling normalization. An important statistical issue with compositional data is the lack of subcompositional coherence, meaning that relative abundances change when data are re-normalized after removing or adding features. While this problem is well documented for small compositions, it has not been investigated in large Omics datasets, which typically contain hundreds or thousands of features and where subcompositions are ubiquitous. Subcompositions arise, for example, when using different reference datasets, sequencing depths or when filtering low-abundant features from the database. In such cases, the most abundant features are preferentially retained, whereas variation between original or full compositions and subcompositions is mainly driven by less abundant features. The standard solution to this problem is the use of logratio transformations, but these complicate interpretations and require handling zeros, which are frequent in Omics data and whose imputation introduces spurious variability.
Methods:
Here, we evaluated subcompositional coherence in five representative Omics datasets: fecal 16S metagenomics, rumen metagenomics (taxonomic and functional levels), liver transcriptomics, and plasma metabolomics, considering both unsupervised and supervised learning contexts. We generated 100 random subcompositions comprising one-third of the original features under an abundance-weighted subcomposition scheme and compared their statistical outputs with those from the full composition.
Results And Discussion:
Raw Omics data showed near-perfect coherence: relative abundances, pairwise correlations and sample distances all exhibited very high (scaled) concordances (≥0.98-0.99). Outputs from commonly used supervised models (linear regression, PLS, random forest, and linear mixed models with a Gaussian kernel) were also highly subcompositionally coherent. We conclude that large Omics datasets expressed as relative abundances are almost subcompositionally coherent when considering a weighted subcomposition scheme, thereby challenging one of the criticisms of using relative data in the Omics field over logratio transformations.
Related Concept Videos
Bioequivalence Data: Statistical Interpretation
Correlation of Experimental Data
For example, a spherical particle moving through a viscous fluid experiences drag. Dimensional analysis shows that the drag force depends on the particle's diameter, velocity, and...
Ordinal Level of Measurement
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks in the...
Nominal Level of Measurement
The data that cannot be measured but can be grouped into categories fall under the nominal level of measurement. Data that is measured using a nominal scale is...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...

