Genotyping, characterization, and imputation of known and novel CYP2A6 structural variants using SNP array data
Alec W R Langlois1,2, Ahmed El-Boraie1,2, Jennie G Pouget2,3
1Department of Pharmacology and Toxicology, University of Toronto, 1 King's College Circle, Toronto, ON, M5S 1A8, Canada.
Journal of Human Genetics
|April 14, 2023
Summary
This study developed a reliable method for genotyping CYP2A6 gene structural variants (SVs), crucial for understanding nicotine metabolism and associated health risks like lung cancer. Imputation of these complex SVs from SNP data proved feasible across ancestries.
Area of Science:
- Pharmacogenomics and Genetic Epidemiology
- Human Genetics and Molecular Biology
Background:
- Cytochrome P450 2A6 (CYP2A6) is key in nicotine metabolism, influencing smoking behavior and lung cancer risk.
- Functional structural variants (SVs) in CYP2A6, including deletions, duplications, and hybrids, complicate accurate genotyping.
- Understanding CYP2A6 SVs is vital for personalized medicine approaches in smoking cessation and cancer prevention.
Purpose of the Study:
- To establish a robust protocol for genotyping complex CYP2A6 structural variants (SVs).
- To functionally characterize known and novel CYP2A6 SVs using phenotypic data.
- To assess the feasibility of imputing CYP2A6 SVs from single nucleotide polymorphism (SNP) array data in diverse populations.
Main Methods:
- Developed a minimal Taqman copy number (CN) assay protocol for CYP2A6 SV genotyping.
- Utilized PCR amplification and Sanger sequencing to characterize a novel CYP2A6-CYP2A7 hybrid SV.
- Phenotyped individuals with SVs using the nicotine metabolite ratio (biomarker of CYP2A6 activity).
- Integrated SV diplotype and SNP array data for phasing and creating ancestry-specific reference panels.
- Employed leave-one-out cross-validation to evaluate CYP2A6 SV imputation accuracy.
Main Results:
- A reliable three-assay protocol for CYP2A6 SV genotyping was successfully developed and validated.
- Replicated known associations between CYP2A6 SVs and metabolic activity.
- Identified and sequenced a novel CYP2A6-CYP2A7 hybrid SV (CYP2A6*53), associated with reduced CYP2A6 activity.
- Demonstrated high feasibility of CYP2A6 SV imputation from SNP array data in European (>70%) and African (>60%) ancestries.
- Achieved low false positive rates (<1%) for imputed CYP2A6 SV alleles.
Conclusions:
- The developed genotyping protocol provides a reliable method for studying complex CYP2A6 structural variants.
- CYP2A6 SV imputation from SNP array data is a feasible and accurate approach for large-scale genetic studies.
- These findings advance the understanding of genetic factors influencing nicotine metabolism and related health outcomes.
Related Concept Videos
Comparing Copy Number Variations and SNPs
17.8K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.8K
Single Nucleotide Polymorphisms-SNPs
15.4K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.4K
Genome-wide Association Studies-GWAS
13.7K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.7K


