用DNA语言模型和图形神经网络对错误变异致病性的疾病特异预测
Mohamed Ghadie1, Sameer Sardaar1, Yannis Trakadis1,2,3
1Research Institute of the McGill University Health Centre (RI-MUHC), Montreal, QC H4A 3J1, Canada.
Bioengineering (Basel, Switzerland)
|October 29, 2025
概括
这项研究引入了一种新的机器学习方法,可以准确预测遗传变异对健康的影响,改善对精准医学不确定意义变异的分类.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 对遗传变异影响的准确预测对于临床遗传学和精准医学至关重要.
- 目前,许多误解变异被归类为具有不确定的意义的变异 (VUS),限制了临床效用.
- 现有的机器学习模型在预测变异病原性方面取得了可变的成功.
研究的目的:
- 开发一种用于疾病特异性预测误解变体致病性的新方法.
- 整合全面的生物医学知识和基因组序列数据,以改善变体解释.
- 减少在临床环境中被归类为VUS的变体数量.
主要方法:
- 使用了一个知识图表,其中有11种相互连接的生物医学实体类型.
- 用BioBERT用于生物医学特征嵌入和DNA语言模型用于变异序列嵌入.
- 实现了两级架构:图形卷积神经网络,然后是神经网络分类器.
- 训练模型通过识别变体和疾病节点之间的边缘来预测疾病特异性变体的致病性.
主要成果:
- 预测平衡精度达到了85.6% (灵敏度:90.5%;净净值:89.8%).
- 证明了整合多样化的生物医学知识和基因组数据的有效性.
- 与错误变体的现有方法相比,展示了较好的分类性能.
结论:
- 开发的方法提供了一个强大的工具,用于准确的疾病特异性变体致病性预测.
- 这种方法有可能显著帮助临床遗传学和推进精准医学.
- 未来的研究可以建立在这个框架上,以进一步完善变异解释和临床决策.
关键词:
克林瓦尔 (ClinVar) 是一个疾病特异性变体解释的解释遗传变异 病原性预测 病原性预测基因组嵌入的基因组嵌入.图形卷积神经网络 (GCN) 是一个神经网络.机器学习 (ML) 是指机器学习.错误的意义变体的变体神经网络分类器神经网络分类器不确定意义的变异 (VUS)更多相关视频
07:15Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
11.3K
08:04Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
1.4K
相关概念视频
Genome-wide Association Studies-GWAS
15.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.3K
Single Nucleotide Polymorphisms-SNPs
17.9K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.9K
