Related Experiment Video
Updated: Feb 11, 2026

Metagenomic Analysis of Silage
Published on: January 13, 2017
Comparison of normalization methods for the analysis of metagenomic gene abundance data
Mariana Buongermino Pereira1, Mikael Wallroth1, Viktor Jonsson1
1Department of Mathematical Sciences, Chalmers University of Technology and University of Gothenburg, SE-412 96, Gothenburg, Sweden.
Background:
In shotgun metagenomics, microbial communities are studied through direct sequencing of DNA without any prior cultivation. By comparing gene abundances estimated from the generated sequencing reads, functional differences between the communities can be identified. However, gene abundance data is affected by high levels of systematic variability, which can greatly reduce the statistical power and introduce false positives. Normalization, which is the process where systematic variability is identified and removed, is therefore a vital part of the data analysis. A wide range of normalization methods for high-dimensional count data has been proposed but their performance on the analysis of shotgun metagenomic data has not been evaluated.
Results:
Here, we present a systematic evaluation of nine normalization methods for gene abundance data. The methods were evaluated through resampling of three comprehensive datasets, creating a realistic setting that preserved the unique characteristics of metagenomic data. Performance was measured in terms of the methods ability to identify differentially abundant genes (DAGs), correctly calculate unbiased p-values and control the false discovery rate (FDR). Our results showed that the choice of normalization method has a large impact on the end results. When the DAGs were asymmetrically present between the experimental conditions, many normalization methods had a reduced true positive rate (TPR) and a high false positive rate (FPR). The methods trimmed mean of M-values (TMM) and relative log expression (RLE) had the overall highest performance and are therefore recommended for the analysis of gene abundance data. For larger sample sizes, CSS also showed satisfactory performance.
Conclusions:
This study emphasizes the importance of selecting a suitable normalization methods in the analysis of data from shotgun metagenomics. Our results also demonstrate that improper methods may result in unacceptably high levels of false positives, which in turn may lead to incorrect or obfuscated biological interpretation.
Related Concept Videos
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
Analysis of Population Pharmacokinetic Data
Renal Drug Clearance: Comparison Between Renal Excretion Methods
Renal clearance is often associated with the renal glomerular filtration rate (GFR), which represents the rate at which plasma is filtered through the glomeruli in the kidney. When drug reabsorption is minimal and there is no active secretion, renal clearance is closely related to the...
Overview of Microsoft Excel as a Data Analysis Tool
The Sense of Self: Reflected Self-Appraisal and Social Comparison
Statistical Methods for Analyzing Epidemiological Data

