一种集体惩罚回归方法用于多祖先多基因风险预测
Jingning Zhang1, Jianan Zhan2, Jin Jin3
1Department of Biostatistics, Johns Hopkins Bloomberg School of Public Health, Baltimore, MD, USA. jingningzhang238@gmail.com.
Nature communications
|April 15, 2024
概括
我们开发了PROSPER,这是一种用于多祖先多基因风险评分 (PRS) 的新方法. 通过整合多样化的全基因组关联研究 (GWAS) 数据,PROSPER改善了少数群体的预测.
科学领域:
- 遗传学 遗传学 是一个
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 先进的多基因风险评分 (PRS) 旨在预测复杂的特征和疾病.
- 目前的PRS方法往往缺乏可转移性,因为培训主要以欧洲祖先数据为主.
- 这限制了它们在多样化的全球人口中的实用性.
研究的目的:
- 引入PROSPER (基于EnSemble of PEnalized Regression模型的多基因风险得分),这是一种用于生成多祖先PRS的新方法.
- 通过利用各种基因组数据,增强PRS对少数群体的预测能力.
- 为大规模,多祖先PRS分析提供计算可扩展的解决方案.
主要方法:
- PROSPER集成了来自多个祖先的全基因组关联研究 (GWAS) 总结统计数据.
- 它采用了 LASSO 和 RIDGE 处罚回归模型的组合.
- 整体方法结合了来自不同惩罚参数的PRS,以获得最佳的性能.
主要成果:
- 在各种遗传架构中,PROSPER在多祖先多基因预测方面取得了实质性的改进.
- 在非洲祖先种群中,与PRS-CSx相比,PROSPER在连续特征中平均增加了70%的R2预测.
- 该方法在现实世界数据集的样本外预测准确度上显示出显著的收益.
结论:
- 在多元祖先之间开发准确和可转移的多基因风险评分方面,PROSPER提供了重大进展.
- 该方法有效地解决了少数群体中现有的PRS的局限性.
- 对于大型基因组数据集和多个种群来说,PROSPER在计算上是高效和可扩展的.
更多相关视频
09:38Generalized Psychophysiological Interaction PPI Analysis of Memory Related Connectivity in Individuals at Genetic Risk for Alzheimer's Disease
Published on: November 14, 2017
14.9K
08:27Large-Scale Multi-Omics Genome-Wide Association Studies Mo-GWAS: Guidelines for Sample Preparation and Normalization
Published on: July 27, 2021
3.7K
相关概念视频
Polygenic Traits
65.8K
When more than one gene is responsible for a given phenotype, the trait is considered polygenic. Human height is a polygenic trait. Studies have uncovered hundreds of loci that influence height, and there are believed to be many more. Due to the high number of genes involved, as well as environmental and nutritional factors, height varies significantly within a given population. The distribution of height forms a bell-shaped curve, with relatively few individuals in the population at the...
65.8K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Genome-wide Association Studies-GWAS
13.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.4K
Heritability
199
Heritability is a statistical concept that measures the degree to which genetic differences among individuals contribute to trait variations within a population. It is a fundamental idea in genetics, often prone to misinterpretation. Heritability is expressed as a percentage, reflecting the proportion of variation in a specific trait across a population that can be linked to genetic differences. However, it's important to understand that heritability does not determine how "genetic"...
199
Pedigree Analysis
84.2K
Overview
84.2K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
