Related Experiment Video
Updated: Jan 18, 2026

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
Finding the Optimal Number of Splits and Repetitions in Double Cross-Fitting Targeted Maximum Likelihood Estimators
Mohammad Ehsanul Karim1,2, Momenul Haque Mondol1,3
1School of Population and Public Health, University of British Columbia, Vancouver, British Columbia, Canada.
For double robust methods like TMLE, using three to five data splits in Double Cross-Fitting (DCF) optimizes statistical performance. More than 25 repetitions do not improve results, and using full data for nuisance estimation is recommended.
Area of Science:
- * Causal inference and statistical learning.
- * Application of machine learning in biostatistics and epidemiology.
Background:
- * Flexible machine learning algorithms in double robust methods (e.g., Targeted Maximum Likelihood Estimator - TMLE) can lead to undercoverage.
- * The Double Cross-Fitting (DCF) procedure allows diverse machine learning estimators but lacks clear guidelines on data splits and repetitions.
Purpose of the Study:
- * To investigate the impact of varying data splits and repetitions in DCF on TMLE estimators.
- * To compare statistical properties of DCF configurations using different split numbers and generalizations.
- * To assess the practical implications of DCF split variations in a real-world health study.
Main Methods:
- * Statistical simulations comparing DCF configurations with varying splits (e.g., 3, 5) and repetitions.
- * Evaluation of two DCF generalizations: equal splits and full data use for nuisance estimation.
- * Application of DCF TMLE to National Health and Nutrition Examination Survey (NHANES) data to study obesity and diabetes risk.
Main Results:
- * Five splits in DCF demonstrated satisfactory bias, variance, and coverage in simulations.
- * DCF TMLE risk difference estimates were consistent across splits in the NHANES analysis, but standard errors increased with more splits in one generalization.
- * Increasing repetitions beyond 25 did not yield performance improvements.
Conclusions:
- * Judicious selection of splits (3-5 recommended) and repetitions is crucial for DCF TMLE methods.
- * Using full data for nuisance estimation in DCF provides more consistent variance estimation.
- * Cautious management of splits in DCF is advised for accurate causal inference with machine learning.
Related Concept Videos
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Expected Frequencies in Goodness-of-Fit Tests
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Distributions to Estimate Population Parameter
Friedman Two-way Analysis of Variance by Ranks
Goodness-of-Fit Test

