Related Experiment Videos
De novo, systemic, deleterious amino acid substitutions are common in large cytoskeleton-related protein coding
Rebecca J Stoll1, Grace R Thompson1, Mohammad D Samy1
1Department of Molecular Medicine, Morsani College of Medicine, University of South Florida, Tampa, FL 33612, USA.
Abstract:
Human mutagenesis is largely random, thus large coding regions, simply on the basis of probability, represent relatively large mutagenesis targets. Thus, we considered the possibility that large cytoskeletal-protein related coding regions (CPCRs), including extra-cellular matrix (ECM) coding regions, would have systemic nucleotide variants that are not present in common SNP databases. Presumably, such variants arose recently in development or in recent, preceding generations. Using matched breast cancer and blood-derived normal datasets from the cancer genome atlas, CPCR single nucleotide variants (SNVs) not present in the All SNPs(142) or 1000 Genomes databases were identified. Using the Protein Variation Effect Analyzer internet-based tool, it was discovered that apparent, systemic mutations (not shared among others in the analysis group) in the CPCRs, represented numerous deleterious amino acid substitutions. However, no such deleterious variants were identified among the (cancer blood-matched) variants shared by other members of the analysis group. These data indicate that private SNVs, which potentially have a medical consequence, occur de novo with significant frequency in the larger, human coding regions that collectively impact the cytoskeleton and ECM.
Insights
Newly identified genetic variants in large human coding regions, particularly those impacting the cytoskeleton and extracellular matrix, occur frequently and may have medical consequences.
Area of Science:
- Genomics
- Molecular Biology
- Human Genetics
Background:
- Human mutagenesis is largely random, making larger coding regions more susceptible to mutations.
- Cytoskeletal-protein related coding regions (CPCRs) and extracellular matrix (ECM) coding regions are substantial genetic targets.
- The study hypothesizes that unique nucleotide variants in these large regions may not be documented in common SNP databases.
Purpose of the Study:
- To investigate the occurrence of systemic nucleotide variants within CPCRs and ECM coding regions.
- To determine if these variants are absent from existing single nucleotide polymorphism (SNP) databases.
- To assess the potential impact of these novel variants on protein function.
Main Methods:
- Utilized matched breast cancer and normal blood-derived datasets from The Cancer Genome Atlas.
- Identified single nucleotide variants (SNVs) in CPCRs not present in the All SNPs(142) or 1000 Genomes databases.
- Employed the Protein Variation Effect Analyzer to predict the functional consequences of identified SNVs.
Main Results:
- Discovered numerous deleterious amino acid substitutions in private CPCR SNVs (not shared among individuals).
- Found no deleterious variants among shared variants within the CPCRs from the analysis group.
- Indicated that private SNVs in large coding regions impacting cytoskeleton and ECM occur frequently de novo.
Conclusions:
- Private SNVs in large coding regions, including those for cytoskeleton and ECM, arise de novo with significant frequency.
- These newly occurring variants have the potential for medical consequences.
- The findings highlight the importance of investigating private genetic variations in large coding regions.