在多重回归中使用简单的随机抽样和排列集抽样方法进行新的调整的缺失值赋值
Juthaphorn Sinsomboonthong1, Saichon Sinsomboonthong2
1Department of Statistics, Faculty of Science, Kasetsart University, Bangkok, Thailand.
PloS one
|March 17, 2025
概括
这项研究评估了多重回归调整的缺失值归算方法. 经调整的回归-多变量比分四分位数1,3 (AR-MRQ1) 方法与排列采样 (RSS) 证明了卓越的效率和准确性,优于简单的随机采样 (SRS).
科学领域:
- 统计 统计 统计 统计
- 数据分析 数据分析
- 回归分析是一种回归分析.
背景情况:
- 缺少数据是统计分析中的一个常见挑战,特别是在多重回归中.
- 有各种各样的归算方法,但它们的效率可以根据采样策略和错误差异而有所不同.
- 调整的归算方法旨在改进处理缺失值的传统技术.
研究的目的:
- 在多重回归分析中比较四种调整后缺失值归算方法的效率.
- 在简单随机抽样 (SRS) 和排列采样 (RSS) 方案下评估这些方法.
- 通过使用不同误差差异的平均平方误差 (MSE) 和平均绝对百分比误差 (MAPE) 度量来评估性能.
主要方法:
- 研究了四种调整的归算方法:R-RQ1,3,AR-CRQ1,3,AR-MRQ1,3和AR-MCRQ1,3.3.
- 这项研究采用了简单随机抽样 (SRS) 和排列采样 (RSS).
- 使用平均平方误差 (MSE) 和平均绝对百分比误差 (MAPE) 量化性能,用于不同的误差差异.
主要成果:
- 用SRS调整的回归-多变量比分四分位数1,3 (AR-MRQ1) 方法显示了小误差差异的最小MSE.
- 调整后回归-多变量链比四分位数1,3 (AR-MCRQ3) 方法在SRS下对大误差差异是最佳的.
- 对于SRS和RSS,AR-MRQ1通常产生最小的MAPE,而AR-MCRQ1对中大误差差异最好.
- 排列采样 (RSS) 估计器始终比简单随机采样 (SRS) 估计器提供更低的MSE和MAPE.
- 总的来说,AR-MCRQ1被确定为多重回归的最有效的归算方法,紧随其后的是AR-MCRQ3.
结论:
- 调整后回归-多变量链比四分位数1,3 (AR-MCRQ1) 方法是多重回归分析中缺失值归算最有效的方法.
- 排列采样 (RSS) 与简单的随机采样 (SRS) 相比,显著提高了归算方法的效率.
- 归算方法和采样策略的选择会影响回归分析结果的准确性.
相关概念视频
Randomized Experiments
6.6K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
6.6K
Random Sampling Method
10.9K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest. Among the various sampling methods used by...
10.9K
Stratified Sampling Method
11.7K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
To choose a stratified sample, divide the population into groups called strata and then take a...
11.7K
Ranks
214
Unlike parametric methods, nonparametric statistics are ideal for nominal and ordinal data, requiring fewer assumptions about the population's nature or distribution. This makes nonparametric methods easier to apply and interpret, as they do not depend on parameters like mean or standard deviation. One common approach in nonparametric analysis is to sort data according to a specific criterion. For instance, we might arrange weather data from hottest to coldest days in a month or rank cities...
214
Systematic Sampling Method
9.9K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
Systematic sampling is one of the simplest methods...
Systematic sampling is one of the simplest methods...
9.9K
Friedman Two-way Analysis of Variance by Ranks
127
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
127


