Related Experiment Video
Updated: Jun 28, 2025

09:35
A Protocol for Using Gene Set Enrichment Analysis to Identify the Appropriate Animal Model for Translational Research
Published on: August 16, 2017
17.8K
bootGSEA: a bootstrap and rank aggregation pipeline for multi-study and multi-omics enrichment analyses
Shamini Hemandhar Kumar1,2, Ines Tapken2,3, Daniela Kuhn3,4
1Institute for Animal Genomics, University of Veterinary Medicine, Foundation, Hannover, Germany.
Frontiers in Bioinformatics
|April 18, 2024
Summary
Gene set enrichment analysis (GSEA) can be unreliable due to database changes. Our bootGSEA pipeline uses bootstrap sampling to assess GSEA robustness, improving biological interpretation and reproducibility in transcriptomics and proteomics.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Gene set enrichment analysis (GSEA) is standard in omics data analysis but faces reproducibility challenges due to dynamic database annotations.
- Changes in gene or protein set compositions can significantly impact biological interpretations derived from GSEA.
Purpose of the Study:
- To introduce bootGSEA, a novel computational pipeline for assessing the robustness and variability of GSEA results.
- To enhance the reliability and reproducibility of GSEA by accounting for changes in set annotations.
Main Methods:
- bootGSEA employs bootstrap sampling to generate multiple GSEA replicates, varying gene/protein inclusion in each.
- Rank aggregation is utilized to combine results from bootstrap replicates and compare them to standard GSEA.
- The pipeline facilitates combining results across different omics levels or multiple studies.
Main Results:
- Application to cancer transcriptomics datasets demonstrated that bootstrap GSEA identifies more robustly enriched gene sets.
- Analysis of paired transcriptomics and proteomics data from a spinal muscular atrophy mouse model showed robust rankings at both omics levels.
- The bootGSEA R-package was developed, offering graphical visualizations and identifying less robust gene/protein sets under varying compositions.
Conclusions:
- Bootstrap-based GSEA enhances the selection of reliable gene sets, mitigating issues from changing database annotations.
- Rank aggregation effectively integrates bootstrap results for robust single-omics or multi-omics findings.
- The bootGSEA pipeline and R-package provide a valuable tool for reproducible and robust GSEA in omics research.

