利用AutoML优化数据集选择,以改善乳腺癌变体的病原性预测
Rahaf M Ahmad1, Noura AlDhaheri1, Mohd Saberi Mohamad1,2,3,4
1Department of Genetics and Genomics, College of Medicine and Health Sciences, United Arab Emirates University, United Arab Emirates.
Computational and structural biotechnology journal
|November 17, 2025
概括
使用自动机器学习 (AutoML) 准确预测乳腺癌 (BC) 变异病原性,选择合适的数据集至关重要. 一个精心策划的数据集,结合了癌症特异性和非癌症数据,在多个AutoML框架中显著提高了预测准确性.
科学领域:
- 基因组医学是基因组医学.
- 计算生物学 计算生物学
- 机器学习在瘤学中
背景情况:
- 乳腺癌 (BC) 是全球癌症死亡的主要原因,受遗传和环境因素的影响.
- 准确预测遗传变异的致病性对于早期检测,风险分层和BC的个性化治疗至关重要.
- 当前的计算工具往往缺乏疾病特异性和变种分类的概括性.
研究的目的:
- 为了对不同自动机器学习 (AutoML) 框架 (TPOT,H2O AutoML,MLJAR) 的性能进行比较,用于乳腺癌变体致病性预测.
- 评估数据集组成对病原性预测模型准确性的影响.
- 为了确定可靠的BC特定变异分类的最佳数据集.
主要方法:
- 使用三个AutoML框架对四个不同的变体数据集进行系统的比较.
- 基于数据集组成的分类性能评估.
- 应用可解释性技术 (SHAP,换重要性,LIME) 来验证模型的透明度和生物相关性.
主要成果:
- 结合癌症特异性和非癌症数据的精选数据集 (Dataset-2) 始终在所有测试的AutoML框架中产生了最高的预测性能.
- 在最佳数据集上,H2O AutoML 实现了 99.99% 的峰值精度,TPOT 和 MLJAR 也表现出强大的性能.
- 特性重要性分析强调了保护得分和病原性指标作为关键预测指标,在各框架之间有很强的共识.
结论:
- 周到的数据集设计,优先考虑与疾病相关的和精选的数据,对于开发基因组医学中准确的机器学习模型至关重要.
- 开发的AutoML框架为乳腺癌变体的临床优先级提供了一个可扩展和可解释的方法.
- 该框架可用于预测其他遗传疾病的致病性,支持精确诊断和个性化瘤学.
相关概念视频
Cancer Survival Analysis
634
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
634
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy
3.5K
3.5K


