Related Experiment Video
Updated: Feb 5, 2026

Spotting Cheetahs: Identifying Individuals by Their Footprints
Published on: May 1, 2016
Development and Validation of an Algorithm for Identifying Patients with Hemophilia A in an Administrative Claims
Jennifer Lyons1, Vibha Desai1, Yaping Xu2
1HealthCore, Wilmington, DE, USA.
Background:
The accuracy with which hemophilia A can be identified in claims databases is unknown.
Objective:
Develop and validate an algorithm using predictive modeling supported by machine learning to identify patients with hemophilia A in an administrative claims database.
Methods:
We first created a screening algorithm using medical and pharmacy claims to identify potential hemophilia A patients in the US HealthCore Integrated Research Database between January 1, 2006 and April 30, 2015. Medical records for a random sample of patients were reviewed to confirm case status. In this validation sample, we used lasso logistic regression with cross-validation to select covariates in claims data and develop a predictive model to estimate the probability of being a confirmed hemophilia A case.
Results:
The screening algorithm identified 2,252 patients and we reviewed medical records for 400 of these patients. The screening algorithm had a positive predictive value (PPV) of 65%. The predictive model identified 18 predictors of being a hemophilia A case or noncase. The strongest predictors of case status included male sex, factor VIII therapy, office visits for hemophilia A, and hospitalizations for hemophilia A. The strongest predictors of noncase status included hospitalizations for reasons other than hemophilia A and factor VIIa therapy. A probability threshold of ≥0.6 resulted in a PPV of 94.7% (95% CI: 92.0-97.5) and sensitivity of 94.4% (95% CI: 91.5-97.2).
Conclusions:
We developed and validated an algorithm to identify hemophilia A cases in an administrative claims database with high sensitivity and high PPV.
Related Concept Videos
Testing a Claim about Mean: Known Population SD
Estimating a population mean requires the samples to be distributed normally. The data should be collected from the randomly selected samples having no sampling bias. The sample size needed to be higher than 30, and most importantly, the population standard deviation should be already known.
In most realistic situations, the population standard deviation is often unknown, but in rare circumstances, when it...
Reliability and Validity
In Vitro Drug Release Testing: Overview, Development and Validation
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Testing a Claim about Mean: Unknown Population SD
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used;...

