基准知识图嵌入模型用于预测寡生组合的基准知识图嵌入模型
Inas Bosch1,2,3,4, Barbara Gravel1,2,3, Alexandre Renaux1,2,3
1Interuniversity Institute of Bioinformatics in Brussels, Université Libre de Bruxelles-Vrije Universiteit Brussel, Boulevard du Triomphe CP263, 1050 Brussels, Belgium.
鉴定罕见疾病的基因对是具有挑战性的. 结构化生物数据和知识图嵌入 (KGE) 显著提高了对寡原性原因的预测准确性.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 鉴定罕见疾病的寡原性原因是由于训练数据有限和预测模型特征选择困难的持续挑战.
- 现有的预测和排名方法在识别复杂遗传疾病的致病基因对的精度有限.
研究的目的:
- 为了研究结构化生物信息的实用性,将其集成到一个异构的知识图中,用于学习遗传表征.
- 为了对最先进的知识图嵌入 (KGE) 模型进行基准测试,以识别涉及寡原性疾病的潜在致病基因对.
主要方法:
- 对各种KGE模型进行了详尽的基准测试,以评估它们在预测致病基因对中的性能.
- 模型使用交叉验证,坚持组和新出现的男性不孕症病例队列进行了评估.
- 特别注意的是,在交叉验证过程中,防止数据在嵌入空间中泄露.
主要成果:
- 基因基因模型在预测致病基因对方面取得了很高的准确性,精度回忆曲线下的面积高达0.93.3.
- 这种表现代表了与以前用于预测寡原性疾病中的基因对的方法相比的显著进步.
- 翻译距离模型 (TransE,MurE,RotatE) 和语义匹配模型 (DisMult,QuatE) 显示出卓越的性能.
结论:
- 结构化生物信息和KGE方法提供了一种强大的方法,用于推进预测涉及寡原性罕见疾病的基因对.
- 仔细的交叉验证至关重要,以避免由于嵌入空间的数据泄漏而导致过度乐观的结果.
- 未来的工作是需要开发的方法,提供解释的鉴定基因对在寡头性疾病的相关性.
更多相关视频
03:37Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
相关概念视频
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Predicting Reaction Outcomes
Behavioral Genetics and Its Designs
The primary methodologies used in behavior genetics include family studies, twin studies, and adoption studies, each providing unique...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Multiple Bar Graph
Each bar or column in the multiple bar graph represents a data value. These graphs are used primarily in interrelating two or more sets of data. The categories of different kinds of data are listed along the horizontal or x-axis, whereas...
