Related Experiment Video
Updated: Aug 20, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Validation of claims-based algorithms to identify non-live birth outcomes
Yanmin Zhu1, Brian T Bateman1,2, Sonia Hernandez-Diaz3
1Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital and Harvard Medical School, Boston, Massachusetts, USA.
Purpose:
Perinatal epidemiology studies using healthcare utilization databases are often restricted to live births, largely due to the lack of established algorithms to identify non-live births. The study objective was to develop and validate claims-based algorithms for the ascertainment of non-live births.
Methods:
Using the Mass General Brigham Research Patient Data Registry 2000-2014, we assembled a cohort of women enrolled in Medicaid with a non-live birth. Based on ≥1 inpatient or ≥2 outpatient diagnosis/procedure codes, we identified and randomly sampled 100 potential stillbirth, spontaneous abortion, and termination cases each. For the secondary definitions, we excluded cases with codes for other pregnancy outcomes within ±5 days of the outcome of interest and relaxed the definitions for spontaneous abortion and termination by allowing cases with one outpatient diagnosis only. Cases were adjudicated based on medical chart review. We estimated the positive predictive value (PPV) for each outcome.
Results:
The PPV was 71.0% (95% CI, 61.1-79.6) for stillbirth; 79.0% (69.7-86.5) for spontaneous abortion, and 93.0% (86.1-97.1) for termination. When excluding cases with adjacent codes for other pregnancy outcomes and further relaxing the definition, the PPV increased to 80.6% (69.5-88.9) for stillbirth, 86.6% (80.5-91.3) for spontaneous abortion and 94.9% (91.1-97.4) for termination. The PPV for the composite outcome using the relaxed definition was 94.4% (92.3-96.1).
Conclusions:
Our findings suggest non-live birth outcomes can be identified in a valid manner in epidemiological studies based on healthcare utilization databases.
Related Concept Videos
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Regression Toward the Mean
Testing a Claim about Mean: Known Population SD
Estimating a population mean requires the samples to be distributed normally. The data should be collected from the randomly selected samples having no sampling bias. The sample size needed to be higher than 30, and most importantly, the population standard deviation should be already known.
In most realistic situations, the population standard deviation is often unknown, but in rare circumstances, when it...
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...

