Unifying the analysis of high-throughput sequencing datasets: characterizing RNA-seq, 16S rRNA gene sequencing and
Andrew D Fernandes1, Jennifer Ns Reid2, Jean M Macklaim2
1, YouKaryote Genomics, London, ON, Canada.
Microbiome
|June 10, 2014
Summary
This study introduces ALDEx2, a robust compositional data analysis tool for high-throughput sequencing. It effectively analyzes diverse datasets like RNA-seq and 16S rRNA gene sequencing, improving accuracy and reducing false positives.
Area of Science:
- Bioinformatics
- Computational Biology
- Statistical Genetics
Background:
- High-throughput sequencing generates complex datasets (e.g., RNA-seq, ChIP-seq, 16S rRNA gene sequencing) with read counts mapped to numerous features.
- Current data analysis methods are experiment-specific and lack cross-applicability.
- Compositional data analysis (CDA) methods, treating data as relative abundances, offer more robust and reproducible analyses.
Purpose of the Study:
- To evaluate the applicability and robustness of the ALDEx2 tool for analyzing diverse high-throughput sequencing data.
- To demonstrate the effectiveness of CDA in unifying analysis across different sequencing experimental designs.
Main Methods:
- Utilized ALDEx2, a Bayesian-based compositional data analysis tool.
- Applied ALDEx2 to three distinct datasets: in vitro selective growth experiment, RNA-seq experiment, and Human Microbiome Project 16S rRNA gene abundance data.
- Assessed ALDEx2's ability to identify differential features, differentially expressed genes, and distinguishing taxa.
Main Results:
- ALDEx2 accurately identified differential abundance in the selective growth experiment.
- It identified a comparable set of differentially expressed genes in RNA-seq data as existing leading tools.
- ALDEx2 successfully distinguished taxa differentiating tongue dorsum and buccal mucosa in the Human Microbiome Project dataset.
- The tool's design minimizes false positives in datasets with many features and few samples.
Conclusions:
- The ALDEx2 R package is a versatile and robust tool for statistical analysis of high-throughput sequencing count data.
- It is suitable for RNA-seq, 16S rRNA gene sequencing, and differential growth experiments, with potential for other similar techniques.
- Compositional data analysis provides a unified and more reliable approach for diverse sequencing data types.
Keywords:
16S rRNA gene sequencingDirichlet distributionMonte Carlo samplingRNA-seqcentered log-ratio transformationcompositional datadifferential abundancehigh-throughput sequencingmicrobiome

