GhostBuster:一个基于深度学习的,文学无偏的基因优先级工具,用于基因注释预测
Giulio Deangeli1, Maria Grazia Spillantini1, Pietro Liò2
1University of Cambridge, Department of Clinical Neurosciences, Clifford Allbutt Building, Hills Road, CB2 0HA Cambridge, UK.
bioRxiv : the preprint server for biology
|July 16, 2025
概括
GhostBuster是一个新的机器学习平台,可以减少基因功能预测中的文献偏见. 它有助于揭示未被充分研究的"幽灵基因"在疾病和生物网络中的作用.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 机器学习 机器学习
背景情况:
- 相当多的人类蛋白质编码基因的特征很差,被称为"幽灵基因".
- 研究文献表现出一种"潮流效应",不成比例地关注注释良好的基因,这引入了偏见.
- 这种文学偏见影响机器学习 (ML) 模型,导致有利于研究良好的基因的预测,并可能高估生物相关性的预测.
研究的目的:
- 开发一个机器学习 (ML) 平台,GhostBuster,旨在预测基因功能,疾病关联和相互作用,同时尽量减少文献偏见.
- 评估有偏见的 (基因本体学) 与无偏见的训练数据集 (LINCS,TCGA,STRING) 对ML模型性能和偏见放大的影响.
主要方法:
- 开发了GhostBuster,一个编码-解码ML平台.
- 在文学偏见数据集上训练的ML模型与在无偏见数据集 (LINCS,TCGA,STRING) 上训练的ML模型进行了比较.
- 评估模型在识别新基因注释和预测基因功能,疾病关联和相互作用方面的有效性.
主要成果:
- 文学偏见的数据集产生了更高的ML指标,但放大了现有的偏见.
- 在无偏见的数据集上训练的模型在识别最近发现的基因注释方面效率高出2-3倍.
- TCGA数据集,文献偏差最小,显示出强大的性能 (ROC-AUC为0.8-0.95).
结论:
- GhostBuster是第一个明确旨在抵消基因注释文献偏见的ML框架.
- 该平台可以预测新型基因功能,完善途径成员资格,并优先考虑基因间GWAS的成功.
- GhostBuster提供了一个强大的工具,用于探索未被充分研究的基因在细胞功能,疾病和分子网络中的作用.
相关概念视频
Genome Annotation and Assembly
19.3K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.3K
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K


