相关实验视频
Updated: Jul 26, 2025

06:35
Basics of Multivariate Analysis in Neuroimaging Data
Published on: July 24, 2010
16.9K
对总结统计数据的处罚回归的多变量扩展,以构建相关特征的多基因风险得分
Meriem Bahda1,2, Jasmin Ricard2, Simon L Girard2,3
1Department of Mathematics and Statistic, Laval University, Québec, QC G1V 0A6, Canada.
HGG advances
|June 19, 2023
概括
多变量拉索改善了基因相关性特征如精神分裂症 (SZ) 和双相情感障碍 (BD) 的多基因风险评分 (PRS) 预测,使用全基因组关联研究总结统计数据. 与单一特征方法相比,这种方法提高了预测准确度.
科学领域:
- 遗传学 是一个遗传学.
- 精神病学是一个精神病学.
- 统计基因组学 统计基因组学
背景情况:
- 人类特征和精神分裂症 (SZ) 和双相情感障碍 (BD) 等疾病之间建立了遗传相关性.
- 从全基因组关联研究 (GWAS) 总结统计数据中结合多个遗传相关性特征的预测因素,可以比单个特征预测因素更好地预测个体特征.
研究的目的:
- 将多个遗传相关性特征的结合预测因素的概念扩展到对GWAS总结统计数据的惩罚回归.
- 引入多变量拉索索姆,一种将SNP对多个特征的影响建模为相关的随机效应的方法,并允许注释依赖的遗传性和遗传共变性.
主要方法:
- 开发了多变量拉索索姆,这是GWAS总结统计的惩罚回归方法,将SNP效应建模为相关的随机效应.
- 纳入基因组注释,以影响SNP对遗传共变性和可遗传性的贡献.
- 使用来自CARTaGENE队列的基因型进行模拟,用于具有类似于SZ和BD的多基因架构的特征.
- 应用多变量拉索素来预测SZ,BD和相关的精神病特征在北克东部SZ和BD相似研究.
主要成果:
- 多变量拉索产生了多基因风险评分 (PRSs),与真正的遗传风险预测因子的相关性更强,与模拟中的现有多特征和单变量方法相比,具有更好的区分能力.
- 在现实应用中,多变量拉索与单变量稀疏PRS相比,与SZ,BD和相关的精神特征有更强的关联,特别是当结合基因组注释时.
结论:
- 多变量拉索索姆证明了使用GWAS总结统计数据改善基因相关性特征的预测的前景.
- 该方法通过利用多特征信息和基因组注释,提供了增强的预测能力.
相关概念视频
Polygenic Traits
66.0K
When more than one gene is responsible for a given phenotype, the trait is considered polygenic. Human height is a polygenic trait. Studies have uncovered hundreds of loci that influence height, and there are believed to be many more. Due to the high number of genes involved, as well as environmental and nutritional factors, height varies significantly within a given population. The distribution of height forms a bell-shaped curve, with relatively few individuals in the population at the...
66.0K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Parametric Survival Analysis: Weibull and Exponential Methods
489
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
489
Biostatistics: Overview
287
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...
287
Correlation and Regression
1.3K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.3K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K

