Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Multiple Regression01:25

Multiple Regression

3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Regression Toward the Mean01:52

Regression Toward the Mean

6.7K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.7K
Regression Analysis01:11

Regression Analysis

7.5K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
7.5K
Correlation and Regression00:53

Correlation and Regression

2.9K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
2.9K
Mechanistic Models: Compartment Models in Individual and Population Analysis01:23

Mechanistic Models: Compartment Models in Individual and Population Analysis

181
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
181
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

8.7K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.7K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Findings and recommendations from a root cause analysis of a clinical trial randomization error.

Journal of clinical and translational science·2026
Same author

Feasibility of Technology-Assisted Lifestyle Self-Monitoring in Older Adults With Type 2 Diabetes: Mixed Methods Pilot Study.

JMIR formative research·2026
Same author

Effect of Sevelamer and B. longum on Insulin Sensitivity in Participants With Obesity: A Randomized Clinical Trial.

Obesity (Silver Spring, Md.)·2026
Same author

Addressing Racial Disparities in a Hispanic Population Through Living Donor Liver Transplantation-A Comparison of 2 Eras.

Clinical transplantation·2026
Same author

Gut Microbiome as a Lifestyle Risk Factor Associated with Prostate Cancer.

European urology focus·2026
Same author

Pragmatic trial design enhances diversity and retention among research participants with Lewy body diseases.

NPJ dementia·2026

Related Experiment Video

Updated: Dec 13, 2025

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
04:35

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach

Published on: July 3, 2020

3.6K

Machine learning outcome regression improves doubly robust estimation of average causal effects.

Byeong Yeob Choi1, Chen-Pin Wang1, Jonathan Gelfond1

  • 1Department of Population Health Sciences, UT Health San Antonio, San Antonio, Texas, USA.

Pharmacoepidemiology and Drug Safety
|July 28, 2020
PubMed
Summary

Machine learning methods, like shrinkage or Super Learner, improve doubly robust estimators for average treatment effect. These techniques enhance robustness against model misspecification, reducing bias and standard error in statistical analyses.

Keywords:
average causal effectcovariate-balancing propensity scoredoubly robust estimationmachine learning techniquesmaximum likelihoodpharmacoepidemiologysimulation

More Related Videos

Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients
07:34

Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients

Published on: August 22, 2018

8.5K
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

7.2K

Related Experiment Videos

Last Updated: Dec 13, 2025

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
04:35

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach

Published on: July 3, 2020

3.6K
Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients
07:34

Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients

Published on: August 22, 2018

8.5K
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

7.2K

Area of Science:

  • Statistics
  • Causal Inference
  • Machine Learning

Background:

  • Doubly robust (DR) estimation offers unbiased average treatment effect (ATE) estimation, provided propensity score (PS) and outcome models are correctly specified.
  • However, DR estimators can be more biased than standard weighting estimators when both PS and outcome models are misspecified.

Purpose of the Study:

  • To evaluate machine learning (ML) methods for estimating conditional potential outcome means.
  • To enhance the robustness of DR estimators against model misspecification by reducing bias and standard error.
  • To compare the performance of various ML techniques within the DR framework.

Main Methods:

  • Assessed ML methods including least squares, tree-based models, generalized additive models, and shrinkage methods for outcome prediction.
  • Employed the Super Learner (SL), an ensemble method combining multiple learners, for estimating conditional means.
  • Conducted simulations across diverse scenarios varying in PS/outcome model complexity and treatment prevalence.

Main Results:

  • Shrinkage methods demonstrated robust performance, yielding low bias and mean squared error in DR estimates, especially with rich models including 2-way covariate interactions.
  • The Super Learner (SL) achieved performance comparable to the best individual method in each simulated scenario.
  • Both shrinkage and SL methods effectively enhanced the accuracy of DR estimators.

Conclusions:

  • Machine learning methods, specifically shrinkage methods with interaction models or the Super Learner, are recommended for improving the accuracy of doubly robust estimators.
  • These advanced methods offer enhanced robustness against model misspecification in causal effect estimation.