高维代因果森林 (hdiCF) 用于使用医疗保健索赔数据识别子组
Tiansheng Wang1, Virginia Pate1, Richard Wyss2
1Department of Epidemiology, Gillings School of Global Public Health, University of North Carolina at Chapel Hill, Chapel Hill, NC.
American journal of epidemiology
|June 13, 2025
概括
一种新的高维方法有效地确定了受益于特定治疗的患者子组,在检测心力衰竭风险的异质治疗效应方面表现优于标准方法. 这种方法有助于个性化医疗,通过精确确定最佳患者群体.
科学领域:
- 药物监督和药物流行病学
- 生物统计学和健康信息学
- 心血管疾病研究研究
背景情况:
- 检测异质治疗效应 (HTE) 对个性化医学至关重要.
- 标准的高维倾向分数 (hdPS) 方法在捕捉复杂的患者特征方面存在局限性.
- 需要新的高维方法来改善HTE检测.
研究的目的:
- 为了比较一种新的高维方法与HTE检测的标准hdPS方法.
- 为了确定具有不同治疗效果的患者亚组,用于-葡萄糖共运输体-2 抑制剂和类似葡萄糖的-1 受体激动剂.
- 评估住院心力衰竭的2年风险差异.
主要方法:
- 使用了一种代因果森林 (iCF) 分组算法.
- 分析了8075名SGLT2抑制剂和7313名GLP-1RA使用者的医疗保险队列.
- 采用了一种新的高维方法 (1个顺序变量/代码) 和标准的hdPS (3个二进制变量/代码).
- 提取了前200个常见的诊断,程序和处方代码.
- 使用逆概率治疗权重评估特定子组的条件平均治疗效应 (CATE).
主要成果:
- 整体人群在心力衰竭住院的风险差异 (aRD) 为-0.4%.
- 新的高维方法将≥2循环利尿剂处方的患者确定为具有最大CATE (aRD:-2.6%) 的子组.
- 标准的hdPS方法确定了患有慢性病的患者 (aRD:-1.7%).
- 灵敏度分析证实了新方法在确定预期的HTE子组方面具有卓越的准确性.
结论:
- 新的高维方法在检测HTE方面表现出比标准hdPS更优异的性能.
- 这种方法准确地识别出患者子组,例如循环利尿剂患者,具有显著的差异性治疗效果.
- 这些发现与先前的临床证据一致,并支持在心血管疾病中使用先进方法进行个性化治疗策略.
相关概念视频
Survival Tree
73
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
73
Comparing the Survival Analysis of Two or More Groups
162
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
162
Statistical Methods for Analyzing Epidemiological Data
330
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
330
Quantifying and Rejecting Outliers: The Grubbs Test
1.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.5K
Cancer Survival Analysis
334
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
334
Causality in Epidemiology
349
Causality or causation is a fundamental concept in epidemiology, vital for understanding the relationships between various factors and health outcomes. Despite its importance, there's no single, universally accepted definition of causality within the discipline. Drawing from a systematic review, causality in epidemiology encompasses several definitions, including production, necessary and sufficient, sufficient-component, counterfactual, and probabilistic models. Each has its strengths and...
349


