通过SNP和年轻的年龄来提高T2D基于机器学习的预测准确性
Cynthia Al Hageh1, Andreas Henschel2,3, Hao Zhou4
1Department of Public Health & Epidemiology, Khalifa University, Abu Dhabi, United Arab Emirates.
Computational and structural biotechnology journal
|July 18, 2025
概括
整合基因组数据适度改进了用于预测2型糖尿病风险的机器学习模型,特别是在年轻人中. 这种方法增强了早期风险识别,并完善了T2D评估.
科学领域:
- 基因组学就是基因组学.
- 机器学习 机器学习
- 糖尿病研究 糖尿病研究
背景情况:
- 2型糖尿病 (T2D) 是一个重大的公共卫生挑战.
- 准确的风险预测对于及时干预至关重要.
- 机器学习 (ML) 模型为改进T2D风险评估提供了潜力.
研究的目的:
- 评估整合临床和基因组数据对ML模型性能对T2D风险预测的影响.
- 用不同的数据组合和年龄组比较模型性能.
主要方法:
- 在发现数据集上训练和测试了六个ML算法.
- 模型使用英国生物银行数据集进行了验证.
- 用单独的临床数据,综合数据和特定年龄的队列来评估表现.
主要成果:
- 基因组数据集成在ML模型中提供了适度的性能提升.
- 家庭病史和高血压等临床因素是关键预测因素.
- 特定的SNP和多基因风险评分 (PRS) 改善了预测,特别是在55岁以下的个体中.
- 模型在联合数据的英国生物库中实现了AUC>91%.
结论:
- 临床因素是强大的T2D预测因素,但基因组数据提供了渐进的改善.
- 整合基因组数据增强了T2D风险预测,特别是在年轻人中早期检测.
- 结合基因组学的多维模型显示出改进T2D风险评估的前景.
更多相关视频
09:37A Phenotyping Regimen for Genetically Modified Mice Used to Study Genes Implicated in Human Diseases of Aging
Published on: July 14, 2016
8.4K
08:04Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
515
相关概念视频
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Single Nucleotide Polymorphisms-SNPs
15.9K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.9K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
iPS Cell Differentiation
2.8K
The ability of induced pluripotent stem cells or iPSCs to differentiate into most body cell types has stimulated repair and regenerative medicine research over the past few decades. iPSC-derived blood cells, hepatocytes, beta islet cells, cardiomyocytes, neurons, and other cell types can repair injuries or regenerate damaged tissue in diseases such as diabetes and neurodegenerative disorders.
2.8K
Genome-wide Association Studies-GWAS
14.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.3K
