基于机器学习的倾向评分估计的评估:对随机试验进行基准观察分析
Kaicheng Wang1, Lindsey Rosman2, Haidong Lu3
1Yale Center for Analytical Sciences, Department of Biostatistics, Yale School of Public Health, New Haven, Connecticut.
medRxiv : the preprint server for health sciences
|June 30, 2025
概括
机器学习倾向评分可能会在心力衰竭研究中引入偏见. 与ML方法相比,传统的后勤回归与专家指导的混选择更好地估计了sacubitril/valsartan的有效性.
科学领域:
- 心血管医学 心血管医学
- 医疗信息学 医疗信息学
- 生物统计学 生物统计学
背景情况:
- 机器学习 (ML) 越来越多地用于倾向性得分估计,以改善共变量平衡和减少偏差.
- 在选择因果推断适当的混因子时,ML的有效性仍然存在争议.
- 心力衰竭的管理通常涉及使用真实世界的数据来比较药物的有效性.
研究的目的:
- 估计萨库比特里尔/瓦尔萨坦与传统疗法对心力衰竭患者全因死亡率的有效性.
- 为了比较使用传统的物流回归与ML方法的倾向性得分估计.
- 将现实世界证据 (RWE) 的发现与 PARADIGM-HF随机对照试验进行比较.
主要方法:
- 美国退伍军人事务部 (2016-2020年) 用可植入心脏转换器除器治疗心力衰竭患者的回顾性队列研究.
- 使用逻辑回归与*先验*混器选择和基于ML的方法 (通用增强模型) 估计了倾向得分.
- 与PARADIGM-HF试验相比,通过全因死亡率,危险比率 (HR) 和风险比率 (RR) 来衡量有效性.
主要成果:
- 逻辑回归与*先验*混器选择产生了结果 (HR=0.93,95% CI 0.61-1.42) 与PARADIGM-HF试验 (HR=0.81,95% CI 0.61-1.06) 密切一致.
- 基于ML的倾向性得分,特别是数据驱动的混者选择,没有超过后勤回归和潜在放大偏差 (HR=0.63,95% CI 0.31-1.30).
- 机器学习方法可能会在复杂的高维数据集中引入过度调整偏差.
结论:
- 对主题的专业知识对于因果推理中的混选择与真实世界的数据至关重要.
- 在这种情况下,传统的逻辑回归与仔细的混者选择可能比ML方法更可靠,以估计治疗的有效性.
- 在没有临床验证的情况下过度依赖数据驱动的方法可能导致心血管研究中的偏见估计.
相关概念视频
Randomized Experiments
7.3K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
7.3K
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
182
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
182
Study Design in Statistics
8.6K
A study design is a set of techniques that allow a researcher to collect and analyze data from different variables defined for a specific research problem. Statistics is commonly for effective study design and more robust experiments,
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
8.6K
Comparing the Survival Analysis of Two or More Groups
299
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
299
Regression Toward the Mean
6.5K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K
Bias in Epidemiological Studies
700
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
700


