Related Experiment Video
Updated: Oct 31, 2025

Large-Scale Multi-Omics Genome-Wide Association Studies Mo-GWAS: Guidelines for Sample Preparation and Normalization
Published on: July 27, 2021
BIGwas: Single-command quality control and association testing for multi-cohort and biobank-scale GWAS/PheWAS data
Jan Christian Kässens1,2, Lars Wienbrandt1, David Ellinghaus1,3
1Institute of Clinical Molecular Biology, Christian-Albrechts-University of Kiel, Rosalind-Franklin-Str. 12, 24105 Kiel, Germany.
BIGwas automates genome-wide association studies (GWAS) for large biobanks, enabling researchers to process 1 million samples efficiently. This scalable pipeline simplifies complex genetic analyses for broader scientific accessibility.
Area of Science:
- Genetics
- Bioinformatics
- Computational Biology
Background:
- Genome-wide association studies (GWAS) and phenome-wide association studies (PheWAS) involving large sample sizes (1 million individuals) from biobanks present significant computational challenges.
- Current methods require substantial time, personnel, and computational resources, with a lack of automated workflow solutions for processing such massive datasets.
Purpose of the Study:
- To present BIGwas, a portable and fully automated pipeline for quality control and association testing of large-scale GWAS data from biobanks.
- To provide a user-friendly solution for researchers, regardless of their bioinformatics expertise or available computational resources.
Main Methods:
- Utilized Nextflow workflow and Singularity software container technology for portability and reproducibility.
- Implemented a fully automated pipeline for quality control and association testing of GWAS data.
- Developed a dynamic parallelization approach for efficient analysis on high-performance compute (HPC) systems.
Main Results:
- BIGwas successfully performs automated quality control and association testing for large-scale GWAS data.
- A single-command GWAS analysis of 974,818 individuals and 92 million genetic markers was completed in approximately 16 days on a small HPC system (7 compute nodes).
- The pipeline is resource-efficient, reproducible, and scalable for local computers or HPC environments.
Conclusions:
- BIGwas empowers researchers with limited bioinformatics knowledge and computational resources to conduct large-scale, multi-cohort GWAS.
- The pipeline facilitates the creation of custom genome-wide phenome-wide association study (PheWAS) resources.
- BIGwas is freely available, promoting wider adoption and accessibility in genetic research.
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Quality Control
Quality control helps track data, visualize trends, and identify variations, making it easier to detect deviations that may affect the accuracy of an analysis. One way to do this is by generating a quality control chart, which...
Comparing the Survival Analysis of Two or More Groups
Test for Homogeneity
Multiple Allele Traits
Statistical Software for Data Analysis and Clinical Trials

