Related Experiment Video
Updated: Aug 7, 2025

05:53
Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
Published on: June 21, 2018
10.2K
SLEMM: million-scale genomic predictions with window-based SNP weighting
Jian Cheng1, Christian Maltecca1, Paul M VanRaden2
1Department of Animal Science, North Carolina State University, Raleigh, NC 27695, United States.
Bioinformatics (Oxford, England)
|March 10, 2023
Summary
We developed SLEMM (Stochastic-Lanczos-Expedited Mixed Models), a new tool for genomic prediction. SLEMM efficiently handles large datasets, offering high accuracy comparable to existing methods for genomic selection.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Genomic data is growing exponentially, making large-scale genomic prediction computationally challenging.
- Accurate genomic prediction is crucial for breeding programs in plants and livestock.
Purpose of the Study:
- To introduce SLEMM (Stochastic-Lanczos-Expedited Mixed Models), a novel software tool designed to overcome computational hurdles in genomic prediction.
- To enhance prediction accuracy through the implementation of SNP weighting within the SLEMM framework.
Main Methods:
- SLEMM utilizes an efficient implementation of the stochastic Lanczos algorithm for Restricted Maximum Likelihood (REML) within a mixed model framework.
- The study involved extensive analyses on seven public datasets across plant and livestock species, comparing SLEMM with other genomic prediction methods.
- Simulations were conducted on datasets with up to 3 million individuals and 1 million SNPs to assess computational performance.
Main Results:
- SLEMM with SNP weighting demonstrated superior predictive ability compared to GCTA's empirical BLUP, BayesR, KAML, and LDAK's BOLT and BayesR models on public datasets.
- In analyses of nine dairy traits from ~300k cows, SLEMM showed comparable prediction accuracies to other methods, with KAML failing to process the data.
- Simulation studies confirmed SLEMM's computational advantage over counterparts for million-scale genomic predictions.
Conclusions:
- SLEMM provides a computationally efficient solution for large-scale genomic predictions, achieving accuracy comparable to established methods like BayesR.
- The software's ability to handle millions of individuals and SNPs makes it a valuable tool for modern genomic selection.
- SLEMM is publicly available, facilitating its adoption in plant and animal breeding research.
Related Concept Videos
Genome-wide Association Studies-GWAS
13.7K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.7K
Comparing Copy Number Variations and SNPs
17.8K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.8K
Single Nucleotide Polymorphisms-SNPs
15.4K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.4K
Polygenic Traits
66.1K
When more than one gene is responsible for a given phenotype, the trait is considered polygenic. Human height is a polygenic trait. Studies have uncovered hundreds of loci that influence height, and there are believed to be many more. Due to the high number of genes involved, as well as environmental and nutritional factors, height varies significantly within a given population. The distribution of height forms a bell-shaped curve, with relatively few individuals in the population at the...
66.1K

