通过预测平均值匹配不完整的顺序和名义数据的推算
Peter C Austin1,2,3, Stef van Buuren4,5
1ICES, Toronto, ON, Canada.
Statistical methods in medical research
|August 18, 2025
概括
与标准回归方法相比,预测平均值匹配为赋值缺失的分类数据提供了更快,更有效的替代方案. 这种方法在多重归算分析中显著减少了计算时间.
科学领域:
- 统计 统计 统计 统计
- 数据科学数据科学数据科学
- 计算统计学 计算统计学
背景情况:
- 通过链式方程进行多变量归算是处理缺失数据的常用方法.
- 分类数据的标准方法包括逻辑回归 (多项式或顺序式).
- 有限的研究存在于预测平均值匹配对分类归因.
研究的目的:
- 为了将预测平均值匹配与逻辑回归方法进行比较,用于归纳分类变量.
- 评估计算负担和统计推断质量.
- 评估各种样本大小和缺失数据率的性能.
主要方法:
- 进行了模拟,以比较归算方法.
- 分析模型包括逻辑和线性回归.
- 变量包括样本大小 (500-5000),缺失率 (5%-50%) 和分类水平 (3-6).
主要成果:
- 预测平均值匹配对多项和顺序逻辑回归进行了有利的表现.
- 这在不同的样本大小和缺失数据百分比中是正确的.
- 预测平均值匹配的计算速度是计算速度的2-6倍.
结论:
- 预测平均值匹配是归类变量的可行方法.
- 它为多重归算提供了大量减少计算时间.
- 这种方法推用于赋值非二元类别变量.
相关概念视频
Ordinal Level of Measurement
25.7K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
25.7K
Nominal Level of Measurement
30.7K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. Not every statistical operation can be used with every set of data. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
The data that cannot be measured but can be grouped into categories fall under the nominal level of measurement. Data that is measured using a nominal...
The data that cannot be measured but can be grouped into categories fall under the nominal level of measurement. Data that is measured using a nominal...
30.7K
Ranks
286
Unlike parametric methods, nonparametric statistics are ideal for nominal and ordinal data, requiring fewer assumptions about the population's nature or distribution. This makes nonparametric methods easier to apply and interpret, as they do not depend on parameters like mean or standard deviation. One common approach in nonparametric analysis is to sort data according to a specific criterion. For instance, we might arrange weather data from hottest to coldest days in a month or rank cities...
286
Wilcoxon Signed-Ranks Test for Matched Pairs
219
The Wilcoxon signed-rank test for matched pairs evaluates the null hypothesis by combining the ranks of differences with their signs. It essentially tests whether the median of the differences in a population of matched pairs is zero. Since the test incorporates more information than the sign test, it generally yields more trustable conclusions. This test also does not require the data to follow a normal distribution, but two conditions must be met for it to be applicable: (1) the data must...
219
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Regression Toward the Mean
6.5K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K


