Related Experiment Video
Updated: Sep 12, 2025

04:35
Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
3.4K
With random regressors, least squares inference is robust to correlated errors with unknown correlation structure.
Zifeng Zhang1, Peng Ding2, Wen Zhou3
1Department of Statistics, Colorado State University, Fort Collins, Colorado 80523, U.S.A.
Biometrika
|August 5, 2025
Summary
Linear regression inference is robust to unknown correlated errors when regressors are random. This finding expands the applicability of linear regression beyond conventional statistical theory, highlighting randomization for robust inference.
Area of Science:
- Statistics
- Econometrics
- Machine Learning
Background:
- Linear regression is a fundamental statistical tool.
- Conventional methods assume fixed regressors and uncorrelated errors, requiring adjustments for known error correlation.
- Existing theory for linear regression breaks down with random regressors and correlated errors.
Purpose of the Study:
- To demonstrate the robustness of linear regression inference with random regressors and unknown correlated errors.
- To challenge existing theoretical limitations in linear regression analysis.
- To explore the impact of error correlation on statistical power.
Main Methods:
- Proving asymptotic normality of t-statistics using novel probabilistic analysis of self-normalized statistics.
- Establishing Berry-Esseen bounds for t-statistics.
- Analyzing the local power of t-tests under weak signal conditions.
Main Results:
- Linear regression inference is robust to unknown correlated errors when regressors are random.
- Asymptotic normality of least squares coefficients does not hold in this regime.
- Error correlation can surprisingly enhance statistical power in the presence of weak signals.
Conclusions:
- Linear regression is applicable in a broader range of scenarios than previously understood.
- Randomization is a valuable technique for ensuring the robustness of statistical inference.
- The findings necessitate a re-evaluation of standard linear regression theoretical assumptions.
More Related Videos
Related Concept Videos
Random and Systematic Errors
12.6K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
12.6K
Correlation and Regression
1.9K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.9K
Multiple Regression
3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K
Residuals and Least-Squares Property
7.8K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.8K
Random Error
1.6K
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
1.6K
Regression Analysis
6.0K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
6.0K

