对比基于森林的随机缺失归算方法,用于倾向性评分分析中的共变量
Yongseok Lee1, Walter L Leite2
1Bureau of Economic and Business Research (BEBR), University of Florida.
Psychological methods
|June 13, 2024
概括
这项研究评估了在倾向性评分分析 (PSA) 中缺少数据的归算方法. 近距离归算 (PI) 和missForest在减少治疗效果估计偏差方面表现出卓越的表现.
科学领域:
- 统计 统计 统计 统计
- 生物统计学 生物统计学
- 观测研究方法 观测研究方法
背景情况:
- 选择偏差是观察性研究中的一个重大挑战.
- 在共变量中缺少数据使得倾向性得分分析 (PSA) 变得复杂.
- 有效处理缺失的数据对于可靠的PSA结果至关重要.
研究的目的:
- 评估使用随机森林的多种归算方法来处理PSA中缺少的共变量.
- 通过链式方程进行多变量归算的性能比较 - 随机森林 (Caliber),近距离归算 (PI) 和missForest.
- 评估归算方法对平均治疗效果估计偏差的影响.
主要方法:
- 使用蒙特卡洛模拟来评估归算方法的性能.
- 评估的方法包括Caliber,近距离归算 (PI) 和missForest.
- 倾向性评分分析应用于现实世界的数据集 (早期儿童纵向研究).
主要成果:
- 近距离归算 (PI) 和missForest在减少平均治疗效果偏差方面表现出卓越的表现.
- 在不同的样本大小和缺失数据机制中,PI和missForest的有效性是一致的.
- 该研究提供了这些方法在评估教育干预的实际演示.
结论:
- 在倾向得分分析中,建议使用近距离归算 (PI) 和missForest来处理缺少的共变量数据.
- 这些方法提供了强大的解决方案,以减轻观察性研究中的选择偏差.
- 对治疗效应的准确估计是通过适当的归算技术来增强缺失数据的精确估计.
相关概念视频
Comparing the Survival Analysis of Two or More Groups
177
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
177
Randomized Experiments
6.9K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
6.9K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Assumptions of Survival Analysis
121
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
121
Goodness-of-Fit Test
3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
121
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
121


