Genomic prediction using subsampling
Alencar Xavier1, Shizhong Xu2, William Muir3
1Department of Agronomy, Purdue University, 915 W. State St., Lilly Hall, West Lafayette, IN, 47907, USA.
BMC Bioinformatics
|March 26, 2017
Summary
Subsampling bootstrap Markov chain significantly reduces computational time for genomic prediction models, with minimal impact on prediction accuracy. This method offers an efficient approach for genetic improvement in plants and animals.
Area of Science:
- Genomics
- Animal Breeding
- Plant Breeding
Background:
- Genome-wide assisted selection is crucial for genetic improvement in plants and animals.
- Whole-genome regression models in a Bayesian framework are standard for genomic prediction.
- Fitting these models with large datasets presents significant computational challenges.
Purpose of the Study:
- To introduce and evaluate the subsampling bootstrap Markov chain method for genomic prediction.
- To assess the impact of subsampling on both prediction accuracy and computational efficiency.
Main Methods:
- The study proposes a subsampling bootstrap Markov chain approach for fitting whole-genome regression models.
- This method involves subsampling observations within each round of a Markov Chain Monte Carlo simulation.
- The impact of subsampling on prediction and computational parameters was evaluated across various datasets.
Main Results:
- An optimal subsampling proportion of approximately 50% with replacement and 33% without replacement was identified.
- Subsampling reduced model fitting time by approximately 50%.
- Losses in predictive properties due to subsampling were negligible, typically less than 1%, with occasional slight improvements observed.
Conclusions:
- The subsampling bootstrap Markov chain algorithm effectively reduces the computational burden of model fitting in genomic prediction.
- Combining subsampling with Gibbs sampling forms an effective ensemble algorithm.
- This method shows potential for enhancing prediction properties while significantly improving computational efficiency.
Related Concept Videos
Prediction Intervals
3.5K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.5K
Downsampling
734
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
734
Predicting Products: Substitution vs. Elimination
15.0K
When a nucleophile and an alkyl halide react, nucleophilic substitution and β-elimination reactions compete to generate products.
The following factors can influence the mechanisms competing against each other:
The following factors can influence the mechanisms competing against each other:
15.0K
End Point Prediction: Gran Plot
1.3K
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
1.3K
Comparing Copy Number Variations and SNPs
18.9K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.9K
Genome-wide Association Studies-GWAS
16.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
16.2K


