在回归分析中,均值的集中是不必要的,而且可能会增加误解系数的风险
1Psychology Department, Gonzaga University, Spokane, WA, United States.
Frontiers in psychology
|July 31, 2025
概括
在回归分析中,以平均值为中心的连续变量对于准确解释主要效应或相互作用是不必要的. 虽然在某些情况下可能有助于理解,但统计软件的进步已经消除了许多先前的担忧.
科学领域:
- 统计 统计 统计 统计
- 量化心理学 量化心理学
- 计量经济学 计量经济学
背景情况:
- 学者越来越多地使用线性建模,特别是涉及连续变量的相互作用.
- 在统计学教科书中,关于平均值中心化的必要性存在重大分歧.
- 对回归系数的误解很常见,特别是在适度回归中.
研究的目的:
- 审查关于统计建模中平均值中心化的建议.
- 为了证明平均值对普通最小方程 (OLS) 回归的中心作用.
- 澄清回归系数的解释,并探索标准化系数的替代方案.
主要方法:
- 对统计教科书关于平均值中心化的建议进行审查.
- 使用连续预测器的OLS回归模型进行两次演示.
- 对数值估计,置信区间和p值的评估,均值集中和不集中.
主要成果:
- 平均中心对估计,置信区间或主要效应,相互作用或OLS回归中的二次数项的p值没有影响.
- 关于平均中心的传统智慧往往是错误的或过时的.
- 标准化回归系数 (β) 有局限性;半局部相关系数 (sr) 有优势.
结论:
- 对具有连续预测器的OLS模型来说,平均中心化不是必需的,尽管它可能有助于解释.
- 计算精度的进步解决了一些关于平均中心化的历史问题.
- 提供了用于有效评估回归模型的实际建议.
更多相关视频
06:33Decomposing the Variance in Reading Comprehension to Reveal the Unique and Common Effects of Language and Decoding
Published on: October 11, 2018
6.9K
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
9.3K
相关概念视频
Regression Toward the Mean
6.5K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K
Statistical Analysis: Overview
7.4K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
7.4K
Central Tendency: Analysis
215
Measures of central tendency are tools used in biostatistics to identify the average or center of a dataset. They offer a single representative value for understanding and summarizing data distribution.
The mean is one such measure, calculated by totaling all values in a dataset and dividing by the number of values. For instance, the mean blood pressure reading (120, 130, 140, 150) would be 135. However, the mean can be affected by extreme values or outliers.
The median, another measure,...
The mean is one such measure, calculated by totaling all values in a dataset and dividing by the number of values. For instance, the mean blood pressure reading (120, 130, 140, 150) would be 135. However, the mean can be affected by extreme values or outliers.
The median, another measure,...
215
Calculating and Interpreting the Linear Correlation Coefficient
6.4K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
6.4K
Trimmed Mean
3.0K
While measuring the mean of a data set, care needs to be taken when associating the mean to its central tendency. The same goes for the arithmetic mean, the geometric mean, or the harmonic mean. This is because the presence of a single outlier data value can significantly affect the mean. That is, the mean is sensitive to fluctuations in the data set.
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
3.0K
Measures of Central Tendency
16.2K
The "center" of a data set is also a way of describing location. The two most widely used measures of the "center" of the data are the mean (average) and the median. The words "mean" and "average" are often used interchangeably. The substitution of one word for the other is common practice. The technical term is "arithmetic mean" and "average" is technically a center location. However, in practice among non-statisticians,...
16.2K
