链接收缩以改善回归模型中相互作用效应的估计
Mark A van de Wiel1, Matteo Amestoy1, Jeroen Hoogland1
1Department of Epidemiology and Data Science, Amsterdam Public Health Research Institute, Amsterdam University Medical Centers, Amsterdam, The Netherlands.
Epidemiologic methods
|July 11, 2024
概括
本研究引入了一种新的统计模型,有效地处理数据中的复杂相互作用,改善预测和变量选择. 该方法为研究人员提供了准确的参数估计和增强的解释性.
科学领域:
- 统计 统计 统计 统计
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
背景情况:
- 高维数据为统计建模带来了挑战,特别是在结合双向交互时.
- 现有的方法通常通过只关注相关的主要效应来简化交互,从而可能限制模型范围.
研究的目的:
- 开发一种估计方法,能够管理由相互作用引起的二次维度增加.
- 创建计算工具来量化变量的重要性,并提高模型的可解释性.
主要方法:
- 提出了一个局部收缩模型,将相互作用效应的收缩与相应的主要效应联系起来.
- 对于Shapley值的新型分析公式是为了快速,个体特定的变量重要性评估而衍生出来的.
主要成果:
- 开发的方法证明了准确的参数估计和竞争力的预测准确性.
- 贝叶斯框架促进了固有的推断和变量选择.
- 对大规模队列数据的实证评估验证了该方法的性能.
结论:
- 链接的局部收缩模型为在流行病学和临床研究中处理相互作用提供了有效的替代方案.
- 该方法提高了参数准确性,预测和变量选择,同时提供了强大的推断和解释.
- 它为预测任务提供了与较少可解释的机器学习算法相比的竞争选择.
更多相关视频
06:52Using Cholesky Decomposition to Explore Individual Differences in Longitudinal Relations between Reading Skills
Published on: September 17, 2019
6.3K
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
9.2K
相关概念视频
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Variation
6.8K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
6.8K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Correlation and Regression
1.2K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.2K
