Related Experiment Video
Updated: Feb 9, 2026

TBase - an Integrated Electronic Health Record and Research Database for Kidney Transplant Recipients
Published on: April 13, 2021
Comparing methods for estimation of heterogeneous treatment effects using observational data from health care
T Wendling1, K Jung2, A Callahan2
1Centre for Health Informatics, Australian Institute of Health Innovation, Macquarie University, Sydney, Australia.
Abstract:
There is growing interest in using routinely collected data from health care databases to study the safety and effectiveness of therapies in "real-world" conditions, as it can provide complementary evidence to that of randomized controlled trials. Causal inference from health care databases is challenging because the data are typically noisy, high dimensional, and most importantly, observational. It requires methods that can estimate heterogeneous treatment effects while controlling for confounding in high dimensions. Bayesian additive regression trees, causal forests, causal boosting, and causal multivariate adaptive regression splines are off-the-shelf methods that have shown good performance for estimation of heterogeneous treatment effects in observational studies of continuous outcomes. However, it is not clear how these methods would perform in health care database studies where outcomes are often binary and rare and data structures are complex. In this study, we evaluate these methods in simulation studies that recapitulate key characteristics of comparative effectiveness studies. We focus on the conditional average effect of a binary treatment on a binary outcome using the conditional risk difference as an estimand. To emulate health care database studies, we propose a simulation design where real covariate and treatment assignment data are used and only outcomes are simulated based on nonparametric models of the real outcomes. We apply this design to 4 published observational studies that used records from 2 major health care databases in the United States. Our results suggest that Bayesian additive regression trees and causal boosting consistently provide low bias in conditional risk difference estimates in the context of health care database studies.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
05:35Author Spotlight: Developing a Point-of-Care Hemoglobin Estimation Method for Anemia Management
Published on: January 19, 2024
Related Concept Videos
Interdisciplinary Care: The Health Care Team-I
Physicians
The physician's primary responsibility is to diagnose illness and direct the medical or surgical treatment of the condition. The authority to admit patients to a healthcare agency or institution and practice care within that setting is granted to physicians by the healthcare agency or institution...
Interdisciplinary Care: The Health Care Team-II
Physical Therapist
A physical therapist (PT) aims to restore function or prevent additional impairment in a patient following an injury or disease. Massage, heat, cold, water, sonar waves, exercises, and electrical stimulation are some treatments used by PTs to treat...
Data Collection by Observations
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
Introduction To Health Care Delivery System
The Institute of Medicine (IOM) advocates for a patient-centered, effective, safe, timely, equitable, and effective healthcare system. The National Priorities...
Traditional Level Of Health Care System
The preventive healthcare service includes tests for screening. Preventive health care services include identifying and reducing disease risk...
Naturalistic Observations