深度学习对预测基因组学的基准解释性:回忆,精度和特征归因变异性
Justin Reynolds1, Chongle Pan1
1School of Computer Science, Gallogly College of Engineering, University of Oklahoma, Norman, Oklahoma, United States of America.
深度神经网络可以预测多基因特征,但识别因果遗传变异是具有挑战性的. 我们的框架显示Saliency归因最好平衡深度学习模型的准确性和稳定性.
科学领域:
- 遗传学和生物信息学 遗传学和生物信息学
- 机器学习在生物学中的应用
背景情况:
- 深度神经网络 (DNN) 可以模拟复杂的多基因特征.
- 鉴定遗传变异的DNN解释性方法的可靠性是不确定的.
- 准确识别遗传变异对于理解特征架构至关重要.
研究的目的:
- 引入和应用一个基准测试框架来量化DNN解释性用于遗传变异识别.
- 评估四个归因算法 (Saliency,梯度SHAP,DeepLIFT,集成梯度) 在识别立身高度预测的因果变异方面的性能.
- 评估SmoothGrad噪声平均化对归因性能和稳定性的影响.
主要方法:
- 开发了一个基准测试框架来衡量归因回忆,精度和稳定性.
- 在英国生物库的基因型数据 (500K+变体,300K参与者) 上训练有素的前神经网络,用于站立高度预测.
- 评估了使用SmoothGrad和没有SmoothGrad的四个归因算法,使用合成spike-in和null decoy变体来评估性能.
主要成果:
- 在最高1%的门上,SmoothGrad的平均归因回忆率提高了~0.16,精度提高了~0.06.
- 各种方法的归因稳定性是可比的,中位数相对标准偏差为0.4-0.5.5.
- Saliency获得了最高的综合分数,在测试方法中显示了回忆,精度和稳定性的最佳平衡.
结论:
- 开发的基准测试框架有效量化了DNN可解释性,用于鉴定遗传变异.
- Saliency 归因提供了一种可靠的方法来识别使用深度学习对多基因特征有贡献的遗传变异.
- 这些发现为在遗传研究中更可靠地应用深度学习提供了基础.
更多相关视频
08:04Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
09:47Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
相关概念视频
Improving Translational Accuracy
Improving Translational Accuracy
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Accuracy and Precision
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
