即使是正确指定的和精确估计的回归模型也可能导致误导
1University of Toronto, Canada.
Accident; analysis and prevention
|October 28, 2023
概括
使用回归模型从观测数据中得出因果结论是具有挑战性的. 即使是完美的模型也可能产生错误的因果洞察力,如果没有充分考虑保持预测变量的现实世界的影响.
科学领域:
- 计量经济学 计量经济学
- 因果推理因果推理
- 统计建模 统计建模
背景情况:
- 回归模型被广泛用于从观测数据中推断因果关系.
- 一个常见的假设是,预测变量可以保持不变,以隔离因果关系.
- 然而,这种假设在现实场景中的有效性往往是可疑的.
研究的目的:
- 批判性地检查使用回归模型从观测数据中得出因果结论的可能性.
- 为了证明即使是完美的回归模型在建立因果关系方面的局限性.
- 为了突出解释因果关系的复杂性,当预测变量保持不变时.
主要方法:
- 用一个思维实验来说明因果推理的挑战.
- 这项研究分析了在回归模型中假设恒定预测变量的含义.
- 以道路安全研究对速度影响的历史案例为例.
主要成果:
- 即使拥有完美的模型和丰富的数据,也可能产生错误的因果关系结论.
- 核心问题在于假设当一个变量被改变时,其他预测变量保持不变.
- 这种假设可能是不可能实现的,或者可能需要未经建模的现实世界的变化.
结论:
- 解释因果推理的回归模型需要仔细考虑保持预测变量的现实世界后果.
- 在回归分析中假设ceteris paribus (所有其他条件均等),当应用于观测数据时可能会有问题.
- 该研究强调了在从单方程回归模型中提取因果关系主张时需要谨慎,特别是在道路安全研究等领域.
相关概念视频
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Survival Tree
88
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
88
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Residual Plots
4.6K
A residual plot is a statistical representation of data used to analyze correlation and regression results. It helps verify the requirements for drawing specific conclusions about correlation and regression. To obtain the residual plot, first, the residual for each data value is calculated, which is simply the vertical distance between the observed and the predicted value obtained from the regression equation.
When the residual values are plotted against the variable x, it is called a residual...
When the residual values are plotted against the variable x, it is called a residual...
4.6K
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K


