Genome-wide detection of human variants that disrupt intronic branchpoints
Peng Zhang1, Quentin Philippot2,3, Weicheng Ren4
1St. Giles Laboratory of Human Genetics of Infectious Diseases, Rockefeller Branch, The Rockefeller University, New York, NY 10065.
Summary
Researchers developed BPHunter, a computational tool to detect disease-causing variants that disrupt branchpoint (BP) sequences in pre-messenger RNA splicing. This tool aids in identifying genetic causes of diseases linked to splicing defects.
Area of Science:
- Molecular Biology
- Genetics
- Bioinformatics
Background:
- Pre-messenger RNA splicing relies on recognizing intronic branchpoint (BP) sequences by spliceosome elements.
- Disruptions in BP sequences by rare variants can lead to disease, but efficient detection in sequencing data was lacking.
- Forty-eight pathogenic variants affecting BP have been identified, highlighting the clinical significance of BP integrity.
Purpose of the Study:
- To develop a genome-wide computational approach for efficiently detecting intronic variants that disrupt branchpoint recognition.
- To establish a comprehensive human genome-wide BP database.
- To identify novel disease-causing variants associated with splicing defects.
Main Methods:
- Integrated existing BP data with new data from DBR1-mutated patients and machine-learning predictions to create a genome-wide BP database.
- Characterized BP and BP-2 features, assessing variation rates and evolutionary conservation.
- Developed and applied BPHunter, a computational tool, to retrospectively and prospectively identify BP variants in sequencing data.
Main Results:
- BPHunter identified 40 of 48 known pathogenic BP variants, providing a strategy for prioritizing candidates.
- BP and BP-2 positions show low variation and high conservation, comparable to exonic regions.
- Prospectively identified novel pathogenic variants in STAT2 (COVID-19) and ITPKB (lymphoma), validated experimentally.
Conclusions:
- BPHunter is an effective genome-wide computational tool for systematically detecting intronic variants disrupting BP recognition.
- The study provides a valuable resource and method for identifying genetic variants underlying splicing-related diseases.
- BPHunter has demonstrated practical utility in identifying novel disease-causing variants, aiding clinical diagnostics and research.
Related Concept Videos
Comparing Copy Number Variations and SNPs
17.9K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.9K
Single Nucleotide Polymorphisms-SNPs
15.6K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.6K
Point and Frameshift Mutations
58
Point mutations are genetic alterations involving the change of a single nucleotide base pair in DNA. Depending on how the alteration affects protein synthesis, they can lead to various consequences.Point mutations fall into the following types:Silent mutations occur when a nucleotide change does not alter the amino acid sequence due to the redundancy of the genetic code. For instance, changing ACC to ACA still encodes threonine, leaving the protein function unaffected. This occurs because...
58
Genome-wide Association Studies-GWAS
14.0K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.0K
Genome Copying Errors
4.3K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.3K
Spontaneous and Induced Mutations
106
Spontaneous mutations arise infrequently during DNA replication due to errors in the process. A key factor behind these errors is tautomeric shifts in nitrogenous bases, where bases transition from keto to enol forms or amino to imino forms. This shift can alter base-pairing rules, leading to mutations. Additionally, reactive oxygen species (ROS) arising from aerobic metabolism can damage DNA, resulting in depurination (loss of a purine base) or depyrimidination (loss of a pyrimidine base).
106


