Related Experiment Video
Updated: Jul 7, 2026

09:35
A Protocol for Using Gene Set Enrichment Analysis to Identify the Appropriate Animal Model for Translational Research
Published on: August 16, 2017
17.9K
Addressing erroneous scale assumptions in microbe and gene set enrichment analysis.
Kyle C McGovern1, Michelle Pistner Nixon2, Justin D Silverman1,2,3,4
1Program in Bioinformatics and Genomics, Pennsylvania State University, State College, Pennsylvania, United States of America.
Plos Computational Biology
|November 20, 2023
Summary
Differential Set Analysis (DSA) of sequence count data can be unreliable due to scale limitations. This study introduces methods to ensure robust and accurate microbial or gene enrichment analysis, even with scale assumption errors.
Area of Science:
- Microbiology
- Bioinformatics
- Genomics
Background:
- Sequence count data, often compositional, lack system scale information.
- Differential Set Analysis (DSA) methods rely on normalization, making implicit scale assumptions.
- Errors in scale assumptions can drastically reduce the accuracy of DSA, with predictive values as low as 9%.
Purpose of the Study:
- To address scale limitations in Differential Set Analysis (DSA) of sequence count data.
- To develop methods for robust and reproducible scale-reliant inference.
- To guide researchers in evaluating the impact of scale limitations on their specific scientific goals.
Main Methods:
- Introduction of a novel sensitivity analysis framework applicable to both simulated and real data.
- Development of a statistical test that controls Type-I error rates despite scale assumption errors.
- Discussion of how scale limitations impact research goals and provision of evaluation tools.
Main Results:
- Commonly used DSA normalization methods make strong, implicit assumptions about unmeasured system scales.
- Small errors in scale assumptions can lead to significantly reduced positive predictive values.
- The proposed sensitivity analysis and statistical test offer robust inference capabilities.
Conclusions:
- Scale limitations in sequence count data analysis can lead to practical inferential errors.
- Rigorous and reproducible scale-reliant inference is achievable through careful application of new methods.
- This work aims to stimulate further research into scale limitations and their impact on biological data analysis.
Related Concept Videos
Genome Copying Errors
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.

