在倾向性得分分析之前对缺失的共变量数据的归算:实践方法的强度的教程和评估
Walter L Leite1, Burak Aydin2, Dee D Cetin-Berber3
1University of Florida, Gainesville, FL, USA.
Evaluation review
|August 1, 2025
概括
多重归算 (MI) 和单一归算 (SI) 方法有效处理缺少的共变量数据用于倾向得分分析 (PSA). 跨MI的表现很好,而SI则足以满足最小的缺失数据,确保在准实验性研究中强有力的偏差减少.
科学领域:
- 流行病学 流行病学
- 生物统计学 生物统计学
- 医疗保健服务研究 医疗服务研究
背景情况:
- 倾向性评分分析 (PSA) 对于减轻准实验研究中的选择偏差至关重要.
- 在PSA之前处理缺少的共变量数据对于准确的倾向性得分估计至关重要.
- 多次归算 (MI) 和单次归算 (SI) 是解决PSA缺失数据的常见方法.
研究的目的:
- 在PSA之前审查MI-within,MI-across和SI方法来处理缺少的共同变量数据.
- 用蒙特卡洛模拟来评估MI-across和SI方法的稳定性.
- 用一个实际的,逐步的例子来说明缺失的数据处理和PSA.
主要方法:
- 蒙特卡洛模拟对连续和分类共变量的MI-across和SI归算策略进行了比较.
- 在模拟条件下,样本大小,共变量数,治疗效果大小,缺失数据机制和缺失数据的百分比各不相同.
- 推算技术包括通过联合建模或通过链式方程 (MICE) 进行多变量推算的MI-across和SI.
主要成果:
- MI-across方法在倾向性得分估计方面表现强.
- 单次归算 (SI) 进行得足够好,特别是缺失数据的百分比较小.
- 一个说明性的例子成功地展示了MI和SI,倾向性得分权重,共变量平衡和治疗效果估计.
结论:
- MI-across是一种强大的方法,用于处理PSA中缺少的共变量数据.
- 当缺失数据百分比较低时,SI可以是一个可行的替代方案.
- 该研究为在现实研究中实施这些方法提供了实用指南.
相关概念视频
Truncation in Survival Analysis
309
Truncation in survival analysis refers to the exclusion of individuals or events from the dataset based on specific criteria related to the time of the event. This exclusion can happen in two primary forms: left truncation and right truncation.
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
309
Censoring Survival Data
241
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
241
Assumptions of Survival Analysis
198
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
198
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
212
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
212
Confounding in Epidemiological Studies
265
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
265
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K


