Related Experiment Video
Updated: May 6, 2026

Large-Scale Multi-Omics Genome-Wide Association Studies Mo-GWAS: Guidelines for Sample Preparation and Normalization
Published on: July 27, 2021
Strategies for developing prediction models from genome-wide association studies
Jincao Wu1, Ruth M Pfeiffer, Mitchell H Gail
1Biostatistics Branch, Division of Cancer Epidemiology and Genetics, National Cancer Institute, Rockville, Maryland, United States of America.
Optimizing genetic risk prediction models for complex diseases requires careful selection of single nucleotide polymorphisms (SNPs). A single-phase procedure for SNP selection and estimation yielded higher accuracy (AUC) than two-phase methods.
Area of Science:
- Genetics and Bioinformatics
- Disease Risk Prediction Modeling
Background:
- Genome-wide association studies (GWASs) identify numerous single nucleotide polymorphisms (SNPs) linked to complex diseases.
- Current risk prediction models using SNPs have limited accuracy, prompting research into improved model building strategies.
- Including a large number of SNPs is hypothesized to enhance predictive performance.
Purpose of the Study:
- To investigate methods for improving the discriminatory accuracy of genetic risk prediction models, measured by the area under the receiver operating characteristic curve (AUC).
- To evaluate different aspects of prediction model building, including SNP selection procedures, data allocation, selection criteria, SNP quantity, and estimation methods.
- To compare single-phase versus two-phase procedures for SNP selection and estimation.
Main Methods:
- Utilized realistic genetic effect size, allele frequency, and linkage disequilibrium (LD) data from GWAS for Crohn's disease and prostate cancer.
- Employed theoretical calculations and simulations to estimate AUC for various model building strategies.
- Assessed SNP selection based on P-value thresholding versus ranking, and compared univariate (marginal) versus multivariate estimation.
Main Results:
- Empirical risk models with 10,000 cases and controls showed significantly lower AUC than theoretically possible.
- The single-phase procedure for SNP selection and estimation outperformed the two-phase procedure in terms of AUC.
- Univariate (marginal) estimation was superior to multivariate estimation, especially in the presence of LD.
Conclusions:
- Initial SNP selection is the most critical step in building accurate genetic risk prediction models.
- For complex diseases and sample sizes of 10,000 or fewer, limiting the number of selected SNPs to tens or hundreds is recommended.
- A single-phase approach for SNP selection and estimation is more effective than a two-phase approach for improving predictive accuracy.
More Related Videos
11:35Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
Published on: August 21, 2016
04:41Mapping Alzheimer's Disease Variants to Their Target Genes Using Computational Analysis of Chromatin Configuration
Published on: January 9, 2020
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Genome Annotation and Assembly
Evolutionary Relationships through Genome Comparisons
Pharmacogenomics: Identification of New Drug Targets