在双重交叉拟合的目标最大概率估计器中找到最佳数量的分割和重复
Mohammad Ehsanul Karim1,2, Momenul Haque Mondol1,3
1School of Population and Public Health, University of British Columbia, Vancouver, British Columbia, Canada.
Pharmaceutical statistics
|September 11, 2025
概括
对于像TMLE这样的双稳定方法,在双交叉拟合 (DCF) 中使用三到五个数据分割优化了统计性能. 超过25次重复并不能改善结果,建议使用完整的数据来估计麻烦.
科学领域:
- * 因果推断和统计学习.
- * 机器学习在生物统计和流行病学中的应用.
背景情况:
- * 灵活的机器学习算法在双强的方法 (例如,有针对性的最大概率估计器 - TMLE) 可以导致覆盖不足.
- *双交叉拟合 (DCF) 程序允许各种机器学习估计器,但缺乏关于数据分割和重复的明确指南.
研究的目的:
- * 调查DCF中变化的数据分割和重复对TMLE估计器的影响.
- * 用不同的分割数和概括来比较DCF配置的统计属性.
- * 评估DCF分裂变异在现实世界健康研究中的实际含义.
主要方法:
- * 统计模拟比较DCF配置与不同的分裂 (例如3,5) 和重复.
- *对两个DCF概括的评估:均等的分割和完整的数据用于麻烦估计.
- *将DCFTMLE应用于国家健康和营养检查调查 (NHANES) 数据,以研究肥胖和糖尿病风险.
主要成果:
- * 在模拟中,DCF中的五个分裂表明了令人满意的偏差,方差和覆盖范围.
- *DCF TMLE风险差异估计在NHANES分析中的分割中是一致的,但标准错误在一个概括中增加了更多的分割.
- *增加25次以上的重复并没有提高表现.
结论:
- *对于DCF TMLE方法来说,谨慎选择分割 (3-5推) 和重复是至关重要的.
- *在DCF中使用完整的数据来估计干扰,可以提供更一致的差异估计.
- *建议谨慎管理DCF中的分裂,以通过机器学习准确地推断因果关系.
相关概念视频
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
1.1K
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
1.1K
Expected Frequencies in Goodness-of-Fit Tests
7.2K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
7.2K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
292
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
292
Distributions to Estimate Population Parameter
5.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
5.0K
Friedman Two-way Analysis of Variance by Ranks
483
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
483
Goodness-of-Fit Test
8.1K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
8.1K


