一个基于SuperLearner的管道,用于开发表型特征的DNA甲基化衍生预测器
Dennis Khodasevich1, Nina Holland2, Lars van der Laan3
1Department of Epidemiology and Population Health, Stanford University School of Medicine, Palo Alto, California, United States of America.
PLoS computational biology
|February 6, 2025
概括
我们使用主成分分析 (PCA) 和整体机器学习开发了新的表观遗传钟. 这些新方法提高了从DNA甲基化数据预测生物年龄和环境暴露的准确性和可靠性.
科学领域:
- 表观遗传学 在表观遗传学中,表观遗传学是指表观遗传学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- DNA甲基化 (DNAm) 反映了环境影响和生物衰老.
- 传统的表观遗传时钟使用CpG位点的惩罚性回归.
- 主要成分分析 (PCA) 显示了培训表观遗传预测器的前景.
研究的目的:
- 使用PCA和整体机器学习开发和评估新的表观遗传预测器.
- 创建一个童年的表观遗传钟,并重建一个成人血液钟.
- 证明这些方法对预测环境暴露的有用性.
主要方法:
- 开发了一个管道来训练三个表观遗传预测器:CPG时钟,PCA时钟和超级学习者PCA时钟 (SL PCA).
- 利用公开可用的DNA甲基化数据集.
- 使用相关系数和中位数绝对误差评估预测器性能.
主要成果:
- 与PCA和CpG时钟相比,SL PCA时钟在数据集中的表型预测准确性得到了改进.
- SL PCA 时钟显示不同DNA甲基化阵列上分析的重复样本之间的一致性更高.
- 表观遗传年龄加速 (EAA) 分析使用SL PCA时钟预测产生了更精确的效果估计.
结论:
- 介绍了一种新的DNAm预测器开发方法,将PCA与集体机器学习 (SuperLearner) 结合起来.
- 这种方法提高了预测器的可靠性,对于纵向研究和复杂的特征预测可能特别有用.
更多相关视频
13:21Comprehensive DNA Methylation Analysis Using a Methyl-CpG-binding Domain Capture-based Method in Chronic Lymphocytic Leukemia Patients
Published on: June 16, 2017
9.9K
14:56Sample Preparation to Bioinformatics Analysis of DNA Methylation: Association Strategy for Obesity and Related Trait Studies
Published on: May 6, 2022
4.4K
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
