Imputation of coding variants in African Americans: better performance using data from the exome sequencing project

Qing Duan1, Eric Yi Liu, Paul L Auer

  • 1Department of Genetics and Department of Computer Science, University of North Carolina, Chapel Hill, NC 27599, USA, Public Health Sciences Division, Fred Hutchinson Cancer Research Center, Seattle, WA 98109, USA, Department of Biostatistics, University of North Carolina, Chapel Hill, NC 27599, USA, Department of Biostatistics and Center for Statistical Genetics, School of Public Health, University of Michigan, Ann Arbor, MI 48109, USA, Renaissance Computing Institute, University of North Carolina, Chapel Hill, NC 27599, USA, Department of Statistics and Department of Genetics, Rutgers University, Piscataway, NJ 08854, USA, Department of Epidemiology, University of North Carolina, Chapel Hill, NC 27599, USA, Department of Epidemiology, University of Washington, Seattle, WA 98195, USA, Division of Epidemiology, Graduate School of Public Health, University of Pittsburgh, Pittsburgh, PA 15261, USA, Department of Epidemiology and Medicine, University of Iowa, Iowa City, IA 52242, Division of Cardiology, George Washington University School of Medicine and Health Sciences, Washington, DC 20037, USA, Department of Preventive Medicine, Keck School of Medicine, University of Southern California/Norris Comprehensive Cancer Center, Los Angeles, CA 90033, USA, Epidemiology Program, University of Hawaii Cancer Center, HI 96813, USA, Division of Genomic Medicine, National Human Genome Research Institute, National Institutes of Health, Bethesda, MD 20892, USA, Department of Molecular Physiology and Biophysics, Center for Human Genetics Research, Vanderbilt University, Nashville, TN 37232, USA, Department of Medicine, Stanford University School of Medicine, Stanford, CA 94305, USA, Division of Endocrinology, Diabetes and Metabolism, Ohio State University, Columbus, OH 43210, USA, Department of Physiology and Biophysics, University of Mississippi Medical Center, Jackson, MS 39216, USA and Department of Genome Sciences, University of Washington, Seattle, WA 98195, USA.

Summary

Using the Exome Sequencing Project reference panel for imputation in African Americans significantly boosts effective sample size for rare coding variants. This approach enhances genetic studies by improving imputation quality compared to the standard 1000 Genomes Project panel.

Related Concept Videos

Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
Principles of Pharmacogenetics: Types of Genetic Variants01:27

Principles of Pharmacogenetics: Types of Genetic Variants

The human genome is over 99.9% identical between individuals, yet genetic differences exist at millions of bases. The human genome contains approximately 3 million variant positions per individual, many of which are heterozygous, contributing to genetic diversity and individual traits. Genetic variations include single-nucleotide polymorphisms (SNPs), insertions, deletions, and copy number variations (CNVs).SNPs, the most common variation, involve single-base changes in DNA. These can be...
Comparing Copy Number Variations and SNPs02:26

Comparing Copy Number Variations and SNPs

Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
Incomplete Dominance01:43

Incomplete Dominance

Gregor Mendel's work (1822 - 1884) was primarily focused on pea plants. Through his initial experiments, he determined that every gene in a diploid cell has two variants called alleles inherited from each parent. He suggested that amongst these two alleles, one allele is dominant in character and the other recessive. The combination of alleles determines the phenotype of a gene in an organism.
Exon Recombination02:32

Exon Recombination

The evolution of new genes is critical for speciation. Exon recombination, also known as exon shuffling or domain shuffling, is an important means of new gene formation. It is observed across vertebrates, invertebrates, and in some plants such as potatoes and sunflowers. During exon recombination, exons from the same or different genes recombine and produce new exon-intron combinations, which might evolve into new genes. 
Exon shuffling follows “splice frame rules.” Each exon has three reading...