Related Experiment Video
Updated: Aug 22, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Population-level and individual-level explainers for propensity score matching in observational studies.
Debashis Ghosh1, Arya Amini2, Bernard L Jones3
1Department of Biostatistics and Informatics, Colorado School of Public Health, Aurora, CO, United States.
Propensity score matching exclusions impact generalizability. Machine learning identifies differences between study and unmatched populations, revealing younger patients are more likely to be included in cancer treatment studies.
Area of Science:
- Observational studies
- Causal inference
- Health services research
Background:
- Propensity score matching is widely used for cancer treatment survival analysis in large databases.
- Exclusion of observations is a common byproduct of propensity score matching schemes.
- These exclusions can be viewed as data-driven eligibility criteria, impacting generalizability.
Purpose of the Study:
- To develop methods for identifying characteristics of unmatched subpopulations in propensity score matching.
- To ascertain the population on which causal effects are estimated.
- To evaluate the representativeness of causal effects using machine learning in cancer registry data.
Main Methods:
- Utilized machine learning, specifically decision trees, to identify inclusion probabilities.
- Applied the method to two datasets from the National Cancer Database.
- Analyzed factors influencing the likelihood of observations being included in the matched sample.
Main Results:
- Decision trees showed younger patients had higher inclusion probabilities (≥0.90) compared to older patients (0.47-0.65) in one dataset.
- Older patients were least likely to be matched, with age being a key factor.
- In the second dataset, both age and surgery status influenced inclusion probability.
Conclusions:
- The study highlights the importance of addressing exclusions in propensity score matching.
- Machine learning can effectively characterize differences between matched and unmatched individuals.
- Complementing propensity score matching with other adjustment methods is recommended.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Bias in Epidemiological Studies
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Analysis of Population Pharmacokinetic Data
Fundamental Attribution Error

