KG2ML:整合知识图和积极的无标签学习来识别与疾病相关的基因
Praveen Kumar1, Vincent T Metzger1, Swastika T Purushotham1
1Department of Internal Medicine, Translational Informatics Division, School of Medicine, University of New Mexico (UNM), Albuquerque, NM, United States.
这项研究介绍了KG2ML,这是一个新的管道,使用正和未标记 (PU) 学习来发现生物医学知识图 (KGs) 中隐藏的疾病基因关联. 该方法成功地确定了潜在的新疾病相关基因,提高了研究能力.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 机器学习 机器学习
背景情况:
- 像DDKG这样的生物医学知识图 (KG) 存储已知的关系,但错过了尚未探索的关联.
- 识别未知的与疾病相关的基因对于推动生物医学研究至关重要.
- 现有的方法耗时,需要高效的计算方法.
研究的目的:
- 开发一种高效的计算方法来识别新型疾病相关基因.
- 克服现有的机器学习管道对知识图分析的局限性.
- 使用先进的机器学习推断以前未知的基因疾病关系.
主要方法:
- 开发了KG2ML (Knowledge Graph to Machine Learning) 管道,使用了一种新的正和无标记 (PU) 学习算法,PULSCAR (正无标记的学习完全随机选择).
- 在KG2ML管道中,从ProteinGraphML中集成基于路径的特征提取.
- 将KG2ML应用于12种疾病,以推断DDKG中不存在的新型疾病相关基因.
主要成果:
- 对于12种疾病,KG2ML确定了排名最高的新型疾病相关基因,其中15种中14种缺乏DDKG的先前明确关联.
- 鉴定的基因显示文献和TINX (目标重要性和新性探索器) 证据的支持.
- 将PULSCAR输入基因作为积极的基因,提高了XGBoost分类性能.
结论:
- 积极和未标记的 (PU) 学习有效地揭示了现有知识图 (KG) 中缺少的疾病基因关联.
- KG2ML管道为生物医学研究提供了一个可扩展和可解释的框架.
- 将KG数据与基于ML的推断集成,通过解决KG的局限性来推动生物医学发现.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
09:58C. elegans Positive Butanone Learning, Short-term, and Long-term Associative Memory Assays
Published on: March 11, 2011
相关概念视频
Chromatin Position Affects Gene Expression
Topologically Associated Domains (TADs)
The 3-dimensional positioning of chromatin in the nucleus influences the...
Velocity and Position by Integral Method
Consider an example to calculate the velocity and position from the acceleration function. A motorboat is traveling at a constant velocity of 5.0 m/s when it starts to decelerate to arrive at the dock. Its acceleration is...
Ogive Graph
Graphing Antiderivatives
Design Example: Identifying the Locations of Monuments in the Field Using Global Positioning System Device
Bar Graph
