重新审视边际性原则:在回归分析中",高阶"术语是否应该总是伴随着"低阶"术语?
Tim P Morris1, Maarten van Smeden2, Tra My Pham1
1MRC Clinical Trials Unit at UCL, London, UK.
Biometrical journal. Biometrische Zeitschrift
|September 30, 2023
概括
统计建模中的边际性原理是取决于上下文的. 它的应用因测量尺度而异,影响模型解释,并可能导致维度的诅咒.
科学领域:
- 统计 统计 统计 统计
- 计量经济学 计量经济学 计量经济学
- 数据分析 数据分析
背景情况:
- 边际性原则建议在统计模型中存在高阶术语时,不要省略低阶术语.
- 这一原则假定有明确的术语等级,把较低级别的术语视为较高级别的术语的"边缘".
研究的目的:
- 检查边际性原则在三个特定的建模环境中的适用性.
- 展示测量尺度如何影响"边际"术语的定义.
- 将这些发现与维度诅咒的概念联系起来.
主要方法:
- 用变量比率分析回归模型的分析.
- 对变量的多项式转换的评估.
- 对干预措施的因数设计的检查.
主要成果:
- 较低级别 ("边际") 项的识别取决于所选择的测量尺度,这可以是任意的.
- 这种规模依赖性使边际性原则的直接应用变得复杂.
- 这些发现为"维度诅咒"现象提供了洞察力.
结论:
- 边缘性原则不是一个普遍的规则,它的实用性取决于具体的环境.
- 分析师在不考虑测量尺度的情况下应用边际性原则时应谨慎行事.
- 重新考虑该原则的应用对于稳健的统计建模是必要的.
相关概念视频
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Correlation and Regression
1.3K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.3K
Regression Analysis
5.8K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.8K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
72
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
72
Mechanistic Models: Compartment Models in Individual and Population Analysis
64
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
64


