Omics data in relative values are almost subcompositionally coherent
Marina Martínez-Álvaro1, Michael Greenacre2, Agustín Blasco1
1Instituto de Ciencia y Tecnología Animal, Universitat Politècnica de València, Valencia, Spain.
Frontiers in Microbiology
|August 7, 2026
Summary
Large omics datasets are nearly subcompositionally coherent when using relative abundances. This finding challenges the necessity of complex logratio transformations for analyzing omics data.
Area of Science:
- Bioinformatics
- Statistical genomics
- Data science
Background:
- Omics data are compositional and often normalized to relative abundances.
- Subcompositional incoherence is a known issue, where relative abundances change upon data subsetting.
- This problem's impact on large omics datasets was previously uninvestigated.
Purpose of the Study:
- To evaluate subcompositional coherence in large omics datasets.
- To compare statistical outputs from full omics compositions versus abundance-weighted subcompositions.
- To assess the implications for using relative abundances versus logratio transformations.
Main Methods:
- Analysis of five diverse omics datasets (metagenomics, transcriptomics, metabolomics).
- Generation of 100 random subcompositions (one-third of original features) using an abundance-weighted scheme.
- Comparison of statistical outputs (unsupervised and supervised learning) between full and subcomposed data.
Main Results:
- Raw omics data exhibited near-perfect subcompositional coherence (concordance ≥0.98-0.99).
- Supervised models (linear regression, PLS, random forest, linear mixed models) also showed high coherence.
- Abundance-weighted subcompositions closely mirrored full composition results.
Conclusions:
- Large omics datasets normalized to relative abundances demonstrate substantial subcompositional coherence.
- The criticism regarding incoherence of relative omics data may be overstated under specific subcomposition schemes.
- This supports the use of relative abundances, potentially simplifying omics data analysis.
Related Concept Videos
Bioequivalence Data: Statistical Interpretation
The statistical interpretation of bioequivalence data is a significant aspect of pharmaceutical research. Bioequivalence refers to the absence of any significant difference in the rate and extent to which the active ingredient in pharmaceutical products becomes available at the site of drug action when administered at the same molar dose under similar conditions. This helps determine if different drug products have similar absorption rates, ensuring their interchangeability.Statistical...
Correlation of Experimental Data
Dimensional analysis simplifies complex physical problems and guides experimental investigations, but it does not provide complete solutions. It identifies the dimensionless groups that influence a phenomenon, but experimental data is needed to establish the specific relationships and validate theoretical predictions.
For example, a spherical particle moving through a viscous fluid experiences drag. Dimensional analysis shows that the drag force depends on the particle's diameter, velocity, and...
For example, a spherical particle moving through a viscous fluid experiences drag. Dimensional analysis shows that the drag force depends on the particle's diameter, velocity, and...
Ordinal Level of Measurement
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks in the...
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks in the...
Nominal Level of Measurement
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. Not every statistical operation can be used with every set of data. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
The data that cannot be measured but can be grouped into categories fall under the nominal level of measurement. Data that is measured using a nominal scale is...
The data that cannot be measured but can be grouped into categories fall under the nominal level of measurement. Data that is measured using a nominal scale is...
How Data are Classified: Numerical Data
Data that are countable or measurable in specific units are called numerical or quantitative data. Quantitative data are always numbers. Quantitative data are the result of counting or measuring the attributes of a population. Amount of money, pulse rate, weight, number of people living in a town, and number of students who opt for statistics are examples of quantitative data.
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Statistical Analysis: Overview
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...

