使用突变部位附近的序列信息对与听力损失相关的变体进行分类和预测
Xiao Liu1, Li Teng1, Jing Sun1
1School of Microelectronics and Communication Engineering, Chongqing University, Chongqing 401331, China.
Computational biology and chemistry
|December 15, 2024
概括
这项研究引入了一种新的计算方法来识别导致听力损失的遗传突变. 该模型准确地预测了与聋相关的突变,有助于更快地诊断听力障碍.
科学领域:
- 遗传学 是一个遗传学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 听力障碍影响全球人口的5%以上,遗传变异是主要原因.
- 识别听力损失的病原性突变是具有挑战性和耗时的.
- 目前的预测工具还没有为特定于聋人的遗传分析进行优化.
研究的目的:
- 开发一种新的计算方法来预测导致聋的突变.
- 为了提高准确性,重点关注侧面突变点的序列区域.
- 帮助诊断和理解遗传性听力障碍.
主要方法:
- 从数据库中提取了与聋相关的突变部位及其附带区域.
- 从七个不同的序列段中计算信息理论特征.
- 训练了五个机器学习算法和一个用于预测的整体模型.
- 使用曲线下的面积 (AUC) 和精度 (ACC) 评估模型性能.
主要成果:
- 该模型实现了0.89的平均AUC和0.85的ACC,用于预测使用250bp侧边区域的致病性聋变异.
- 证明了与聋相关的突变部位的高识别率.
- 成功地应用了一种整体方法来评分和排名不确定的意义 (VUS) 的变体,以确定与聋的潜在关联.
结论:
- 开发的机器学习模型提供了一种有效和高效的方法来识别与聋相关的突变.
- 这种方法可以加速遗传性听力损失的诊断.
- 预测VUS有助于进一步调查潜在的聋相关遗传因素.
更多相关视频
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
33.6K
11:35Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
Published on: August 21, 2016
12.9K
相关概念视频
Comparing Copy Number Variations and SNPs
17.2K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.2K
Genome-wide Association Studies-GWAS
12.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.4K
Signal Sequences and Sorting Receptors
5.2K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
5.2K
Single Nucleotide Polymorphisms-SNPs
14.0K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
14.0K
