Related Experiment Video
Updated: May 20, 2026

Three Differential Expression Analysis Methods for RNA Sequencing: limma, EdgeR, DESeq2
Published on: September 18, 2021
Technical and biological variance structure in mRNA-Seq data: life in the real world
Ann L Oberg1, Brian M Bot, Diane E Grill
1Division of Biomedical Statistics and Informatics, Department of Health Sciences Research, Mayo Clinic, 200 1st St SW, Rochester, MN 55905, USA. oberg.ann@mayo.edu
RNA sequencing count data analysis benefits from using Negative Binomial distribution models. These models better capture biological variability and improve the identification of true biological signals in gene expression studies.
Area of Science:
- Genomics
- Bioinformatics
- Statistical Modeling
Background:
- Next-generation sequencing generates mRNA expression data as gene or exon counts.
- Traditionally, count data assumed a Poisson distribution (variance equals mean).
- The Negative Binomial distribution is often used for count data, accommodating over-dispersion (variance > mean).
Purpose of the Study:
- To evaluate the suitability of Poisson and Negative Binomial distributions for modeling mRNA sequencing count data.
- To assess the impact of over-dispersion and experimental effects on statistical modeling.
- To guide the development of analytical strategies for accurate variance modeling in gene expression studies.
Main Methods:
- Analysis of mRNA sequencing data from 25 subjects.
- Comparison of Poisson, over-dispersed Poisson, and Negative Binomial models.
- Assessment of goodness-of-fit and over-fitting for different models.
- Inclusion of experimental effects in modeling.
Main Results:
- Technical variation in mRNA-Seq data generally followed a Poisson distribution.
- Biological variability exhibited over-dispersion relative to the Poisson model.
- A quadratic mean-variance relationship indicated a Negative Binomial distribution.
- Over-dispersed Poisson and Negative Binomial models showed improved goodness-of-fit over the standard Poisson model, with some evidence of over-fitting.
- Modeling experimental effects improved goodness-of-fit for high-variance genes but exacerbated over-fitting.
Conclusions:
- The Negative Binomial distribution provides a better fit for mRNA-Seq count data due to biological over-dispersion.
- Accurate variance modeling is crucial for reliable statistical analysis and sample size determination.
- Improved analytical strategies will enhance the identification of true biological signals for a better understanding of biological systems.
Related Concept Videos
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Ribosome Profiling
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique helps...
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Genetic Variation
Genes exist in different versions called alleles, which...
Evolutionary Relationships through Genome Comparisons
Leaky Scanning

