使用随机回归器,最小方程推断对与未知相关性结构相关的相关错误具有稳定性
Zifeng Zhang1, Peng Ding2, Wen Zhou3
1Department of Statistics, Colorado State University, Fort Collins, Colorado 80523, U.S.A.
Biometrika
|August 5, 2025
概括
当回归者是随机的时,线性回归推理对未知的相关错误是可靠的. 这一发现扩大了线性回归的适用性,超越了传统的统计理论,突出了可靠推断的随机化.
科学领域:
- 统计 统计 统计 统计
- 计量经济学 计量经济学
- 机器学习 机器学习
背景情况:
- 线性回归是一种基本的统计工具.
- 传统方法假设固定的回归和非相关的错误,需要对已知的错误相关性进行调整.
- 现有的线性回归理论与随机回归器和相关错误崩.
研究的目的:
- 用随机回归器和未知的相关错误证明线性回归推理的稳定性.
- 挑战线性回归分析现有的理论局限性.
- 探索错误相关性对统计功率的影响.
主要方法:
- 使用自我规范化统计数据的新型概率分析证明t统计的非对称正常性.
- 建立 Berry-Esseen 对 t 统计的界限.
- 在弱信号条件下分析t测试的局部功率.
主要成果:
- 当回归者是随机的时,线性回归推理对未知的相关错误是可靠的.
- 最小方位系数的非对称正常性在这种制度中不成立.
- 错误相关性可以在存在弱信号的情况下惊人地提高统计能力.
结论:
- 线性回归适用于比以前理解的更广泛的场景.
- 随机化是一种有价值的技术,可以确保统计推理的稳定性.
- 这些发现需要重新评估标准线性回归理论假设.
相关概念视频
Random and Systematic Errors
12.6K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
12.6K
Correlation and Regression
1.9K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.9K
Multiple Regression
3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K
Residuals and Least-Squares Property
7.8K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.8K
Random Error
1.6K
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
1.6K
Regression Analysis
6.0K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
6.0K


