伪随机数生成器对使用机器学习获得的平均治疗效果估计的影响
Ashley I Naimi1, Ya-Hui Yu1, Lisa M Bodnar2
1From the Department of Epidemiology, Emory University.
Epidemiology (Cambridge, Mass.)
|August 16, 2024
概括
研究中的机器学习结果可以根据随机数发生器种子而有所不同. 这种变异性影响风险差异的估计,敦促谨慎解释依赖于这种种子的发现.
科学领域:
- 流行病学 流行病学
- 医疗保健中的机器学习
- 统计建模 统计建模
背景情况:
- 机器学习 (ML) 用于暴露效应估计,在研究结果和伪随机数生成器 (PRNG) 种子值之间产生依赖性.
- 这种种子依赖性为实证研究结果引入了潜在的变异性.
研究的目的:
- 评估不同PRNG种子值对风险差异估计的影响.
- 为了检查水果/蔬菜消费与产前风险之间的关联,风险差异的可变性.
主要方法:
- 利用了来自10,038名孕妇和10%的子样本 (N=1004) 的数据.
- 使用一个增强的反向概率加权估计器与两个超级学习算法 (简单和复杂).
- 在5000个不同的种子值中评估了风险差异,标准误差和P值.
主要成果:
- 观察到风险差异估计的显著变化,受堆叠算法的影响.
- 风险差异的四分位数间范围宽度在算法和样本大小之间显著不同.
- 风险差异分布的中位数因样本大小和算法复杂性而有所不同.
结论:
- 这些发现突出了关于"p-hacking"的担忧,以及在经验研究中需要先进的证据值.
- 取决于PRNG种子值的结果需要仔细解释.
- 强调了ML驱动的流行病学研究中透明度和可重复性的重要性.
更多相关视频
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
7.5K
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
8.2K
相关概念视频
Randomized Experiments
6.8K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
6.8K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Group Design
8.9K
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between...
8.9K
Random Error
849
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
849
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
45
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
45
Mechanistic Models: Compartment Models in Individual and Population Analysis
33
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
33
