机器学习和双重可靠的估计器在结合高维代理来减少剩余混方面有多有效?
Mohammad Ehsanul Karim1,2, Yang Lei3
1School of Population and Public Health, University of British Columbia, Vancouver, British Columbia, Canada.
Pharmacoepidemiology and drug safety
|May 14, 2025
概括
高维代理可以改善观测研究中的混调整. 使用代理的标准方法是稳健的,而TMLE中的复杂机器学习可能会减少覆盖范围,需要仔细配置.
科学领域:
- 流行病学 流行病学
- 生物统计学 生物统计学
- 机器学习在健康研究中的应用
背景情况:
- 在高维观测研究中,残留混是一个主要的挑战.
- 高维代理调整方法,如hdps,使用代理用于未测量的混因子.
- 机器学习和双倍强大的估计器已经集成到hdps扩展中,但它们的比较性能尚不清楚.
研究的目的:
- 为了评估标准方法的性能,超级学习者 (SL),目标最大概率估计 (TMLE) 和双交叉适合TMLE (DC-TMLE) 在混调整中.
- 使用不同的机器学习学习器配置,在不同的暴露和结果流行下比较这些方法.
- 评估高维代理和学习者复杂性对偏差,覆盖和可变性的影响.
主要方法:
- 进行了等离子模拟以评估方法性能.
- 评估的标准方法,SL,TMLE和DC-TMLE.
- 在三个学习者图书馆大小中比较性能:1,3和4个学习者 (包括后勤回归,MARS,LASSO和XGBoost).
主要成果:
- 没有代理的方法显示了最高的偏差和最低的覆盖率.
- 采用高维代理的标准方法显示出具有低偏差和良好的覆盖率的强大性能.
- TMLE和DC-TMLE减少了偏差,但覆盖范围更差,特别是在更大,更复杂的学习者图书馆.
- DC-TMLE在高维环境中表现不佳,与非Donsker学习者一起,表明不稳定.
结论:
- 高维代理对于标准方法中有效的混调整至关重要.
- 在SL和TMLE中定制机器学习学习者配置对于可靠的混调整至关重要.
- 仔细选择学习者至关重要,以避免不稳定,并确保复杂的观测研究的准确结果.
相关概念视频
Strategies for Assessing and Addressing Confounding
73
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
73
Confounding in Epidemiological Studies
117
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
117
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
34
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
34
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Residuals and Least-Squares Property
7.2K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.2K


