Efficient variance components analysis across millions of genomes.
Ali Pazokitoroudi1, Yue Wu1, Kathryn S Burch2
1Department of Computer Science, UCLA, Los Angeles, CA, 90095, USA.
Nature Communications
|August 13, 2020
Summary
We developed a fast and accurate method for variance components analysis in large genetic datasets. This approach reveals insights into trait heritability and the impact of genetic variation, including evidence of negative selection.
Area of Science:
- Genetics
- Bioinformatics
- Statistical Genetics
Background:
- Variance components analysis is crucial for understanding complex traits.
- Current methods struggle with large genetic datasets.
- Efficient analysis is needed for large-scale genomic studies.
Purpose of the Study:
- To present an accurate and efficient method for variance components analysis.
- To enable analysis of large-scale genetic variation data.
- To estimate and partition SNP-heritability for complex traits.
Main Methods:
- Developed a scalable method for variance components analysis.
- Applied the method to analyze 22 traits in 300,000 individuals.
- Utilized genotypes from approximately 8 million common and low-frequency SNPs.
- Partitioned heritability across 28 functional annotations.
Main Results:
- The method efficiently estimates numerous variance components on large datasets.
- Observed increasing per-allele squared effect sizes with decreasing minor allele frequency (MAF) and linkage disequilibrium (LD).
- Found heritability enrichment in FANTOM5 enhancers for several immune and autoimmune disorders.
Conclusions:
- The new method significantly improves the scalability of variance components analysis.
- Findings suggest negative selection acting on common and low-frequency variants.
- Heritability is enriched in specific functional elements like enhancers for certain traits.
Related Concept Videos
Comparing Copy Number Variations and SNPs
18.4K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.4K
Evolutionary Relationships through Genome Comparisons
6.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.7K
Genetic Variation
1.1K
Genetic variation is the diversity in DNA sequences found among individuals of the same species. This diversity is crucial for a species' survival because it helps organisms adapt to environmental changes. Genetic variation begins with fertilization, where an egg and sperm cell merge. Each of these cells carries 23 chromosomes, up to 46 in the fertilized egg. Chromosomes are long DNA strands that contain genes, the basic units of heredity.
Genes exist in different versions called alleles,...
Genes exist in different versions called alleles,...
1.1K
Variability: Analysis
360
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
360
Variance
11.6K
The deviations show how spread out the data are about the mean. A positive deviation occurs when the data value exceeds the mean, whereas a negative deviation occurs when the data value is less than the mean. If the deviations are added, the sum is always zero. So one cannot simply add the deviations to get the data spread. By squaring the deviations, the numbers are made positive; thus, their sum will also be positive.
The standard deviation measures the spread in the same units as the data....
The standard deviation measures the spread in the same units as the data....
11.6K
Hardy-Weinberg Principle
75.6K
Diploid organisms have two alleles of each gene, one from each parent, in their somatic cells. Therefore, each individual contributes two alleles to the gene pool of the population. The gene pool of a population is the sum of every allele of all genes within that population and has some degree of variation. Genetic variation is typically expressed as a relative frequency, which is the percentage of the total population that has a given allele, genotype or phenotype.
75.6K


