Related Experiment Video
Updated: Jul 12, 2025

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
Even correctly specified and well-estimated regression models can mislead
1University of Toronto, Canada.
Abstract:
It should be possible to draw causal conclusions from happenstance data. However, there are many well-known reasons for doubting the causal interpretation of single equation regression models based on such data. Still, hope springs eternal. The hope is founded on the belief that if the function linking the response variable to the predictor variables was known and its parameters estimated from plentiful data then one could predict what change in the response variable is caused by a change in a predictor variable. But what if this foundational belief was incorrect? I use a thought experiment to show even perfect models can lead to incorrect conclusions. The problem is that to say what change in the response variable is caused by a change in a predictor variable one must assume that all the other predictor variables remain unchanged. This may not be possible or may require changes to reality that are outside of the model, changes that almost certainly will not exist. To interpret the estimated model equation correctly one must trace all real-world consequences of holding the predictor variables constant. This is not easy to do. The history of regression-based research about the road safety effect of speed supports my case.
Related Concept Videos
Regression Toward the Mean
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Survival Tree
Building a Survival Tree
Constructing a...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Residual Plots
When the residual values are plotted against the variable x, it is called a residual...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...

