特征归因增强了一些特征的非线性遗传预测
Ruoyu He1,2, Jinwen Fu1,2, Jingchen Ren1,2
1Division of Biostatistics and Health Data Science, School of Public Health, University of Minnesota, Minneapolis, MN 55414, USA.
Genetics
|September 10, 2024
概括
这项研究开发了一种方法,使用来自生物库的遗传数据来归咎缺失的表型. 这种方法提高了对某些特征的非线性多基因 (风险) 评分的准确性.
科学领域:
- 遗传学 是一个遗传学.
- 生物信息学是一种生物信息学.
- 生物医学研究生物医学研究
背景情况:
- 生物库包含大量的遗传和表型数据,对生物医学研究至关重要.
- 缺少的表型数据是有效利用生物库资源的一个主要限制.
- 开发精确的遗传预测模型,如多基因 (风险) 评分 (PGS),对于个性化医学至关重要.
研究的目的:
- 使用全基因组关联研究 (GWAS) 总结数据,在大型基因型数据集中归因缺失的表型.
- 将假定的表型与现有的完整数据集集集成,以构建改进的非线性多基因 (风险) 评分模型.
- 评估与传统方法相比,与归入的表型训练的非线性模型的预测性能.
主要方法:
- 利用大型基因型数据集 (例如,英国生物库),缺乏特定的表型.
- 利用GWAS总结统计数据来归因于个人缺失的表型数据.
- 训练有素的非线性预测模型使用观察和归算的表型数据.
- 开发了一个集合模型来整合来自多个非线性模型的预测.
- 通过计算数据与观察数据训练的模型的R平方值进行了比较.
主要成果:
- 用指定的表型训练的集体模型比只使用小,完整的观察数据集的模型实现了更高的预测准确性 (R2).
- 对于七种特征中的两种,使用归算表型训练的非线性模型比直接使用归算表型作为多基因 (风险) 评分的非线性模型表现更好.
- 对于剩余的五个特征,当直接使用归算特征作为PGS时,没有观察到显著的改善.
- 证明了归算和非线性建模的潜力,以提高遗传预测的准确性.
结论:
- 从大型基因型数据集中归纳缺失的表型是加强生物医学研究的可行策略.
- 通过先进的建模来计算非线性遗传关系,可以提高特定特征的多基因 (风险) 评分的准确性.
- 这种方法提供了一种有希望的方法来克服生物库中的数据限制,并推进遗传预测.
相关概念视频
Truncation in Survival Analysis
179
Truncation in survival analysis refers to the exclusion of individuals or events from the dataset based on specific criteria related to the time of the event. This exclusion can happen in two primary forms: left truncation and right truncation.
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
179
Improving Translational Accuracy
2.5K
2.5K
Heritability
195
Heritability is a statistical concept that measures the degree to which genetic differences among individuals contribute to trait variations within a population. It is a fundamental idea in genetics, often prone to misinterpretation. Heritability is expressed as a percentage, reflecting the proportion of variation in a specific trait across a population that can be linked to genetic differences. However, it's important to understand that heritability does not determine how "genetic"...
195
Polygenic Traits
65.6K
When more than one gene is responsible for a given phenotype, the trait is considered polygenic. Human height is a polygenic trait. Studies have uncovered hundreds of loci that influence height, and there are believed to be many more. Due to the high number of genes involved, as well as environmental and nutritional factors, height varies significantly within a given population. The distribution of height forms a bell-shaped curve, with relatively few individuals in the population at the...
65.6K
Genetic Drift
39.6K
Natural selection—probably the most well-known evolutionary mechanism—increases the prevalence of traits that enhance survival and reproduction. However, evolution does not merely propagate favorable traits, nor does it always benefit populations.
39.6K
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K


