概括
一种新的无监督方法,UNSemblePRS,结合了预先训练的多基因风险评分 (PRS) 模型,而不需要目标人口数据. 这种方法提高了遗传风险预测准确性,用于现实世界的应用.
科学领域:
- 遗传学 遗传学 是一个
- 机器学习 机器学习
- 生物信息学是一种生物信息学.
背景情况:
- 预先训练的多基因风险评分 (PRS) 模型越来越多地可用于现实世界.
- 挑战包括PRS模型的可转移性,数据异质性和目标人群中缺乏表型数据.
- 现有的集合方法通常需要目标人群数据或全基因组关联研究 (GWAS) 来进行优化.
研究的目的:
- 开发一个无监督的集体学习框架,将预先训练的PRS模型结合起来.
- 为了能够准确地预测遗传风险,而不需要目标人群的表型数据.
- 为了促进PRS集成到现实世界的应用程序.
主要方法:
- 开发了UNSupervised enSemble PRS (UNSemblePRS),这是一个不受监督的集体框架.
- 基于预测一致性的聚合预训练PRS模型.
- 在"我们所有人"数据库中使用连续和二进制特征评估性能.
主要成果:
- UNSemblePRS在各种人群中展示了可扩展性和强大的性能.
- 该框架成功地结合了没有表型数据的PRS模型.
- 在现实环境中实现了准确的遗传风险预测.
结论:
- UNSemblePRS是一个可访问的工具,用于将各种PRS模型集成到临床实践中.
- 无监督方法克服了传统监督方法的局限性.
- 随着PRS模型的可用性扩大,提供了广泛的适用性.
更多相关视频
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.4K
09:38Generalized Psychophysiological Interaction PPI Analysis of Memory Related Connectivity in Individuals at Genetic Risk for Alzheimer's Disease
Published on: November 14, 2017
14.8K
相关概念视频
Polygenic Traits
64.4K
When more than one gene is responsible for a given phenotype, the trait is considered polygenic. Human height is a polygenic trait. Studies have uncovered hundreds of loci that influence height, and there are believed to be many more. Due to the high number of genes involved, as well as environmental and nutritional factors, height varies significantly within a given population. The distribution of height forms a bell-shaped curve, with relatively few individuals in the population at the...
64.4K
Genome-wide Association Studies-GWAS
12.1K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.1K
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Multiple Allele Traits
33.8K
The Concept of Multiple Allelism
33.8K
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K
