Related Experiment Video
Updated: Jan 4, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Quantifying outcome misclassification in multi-database studies: The case study of pertussis in the ADVANCE project
Rosa Gini1, Caitlin N Dodd2, Kaatje Bollaerts3
1Agenzia regionale di sanità della Toscana, Osservatorio di epidemiologia, Florence, Italy.
This study developed a method to quantify misclassification in vaccine safety data. Comparing algorithms for Bordetella pertussis (BorPer) cases in European databases revealed potential biases, aiding accurate vaccine benefit-risk monitoring.
Area of Science:
- Pharmacovigilance and Pharmacoepidemiology
- Vaccine Safety Monitoring
- Health Informatics and Database Analysis
Background:
- The Accelerated Development of Vaccine Benefit-risk Collaboration in Europe (ADVANCE) project aims to enhance vaccine safety monitoring using European healthcare databases.
- Accurate estimation of vaccine benefit-risk relies on precise identification of adverse events, which can be challenging due to event misclassification in large databases.
- Manual chart review is often infeasible for rapid monitoring, necessitating automated strategies to quantify misclassification.
Purpose of the Study:
- To develop and test a strategy for quantifying event misclassification in vaccine safety data when manual review is not feasible.
- To evaluate different algorithms for identifying Bordetella pertussis (BorPer) cases across multiple European healthcare databases as a case study.
- To estimate the validity of these algorithms, including positive predictive value (PPV) and sensitivity, by comparing them with external surveillance data.
Main Methods:
- Utilized four primary care (PC) databases (BIFAP, THIN, RCGP RSC, PEDIANET) and one PC/hospital database (SIDIAP) across Europe.
- Defined BorPer algorithms based on healthcare setting, data domain (diagnoses, drugs, lab tests), and concept sets (specific/unspecified pertussis).
- Estimated BorPer incidence rates (IRs) in children (0-14 years) and approximated validity indices (PPV, sensitivity) using novel formulas in SIDIAP.
Main Results:
- Incidence rates varied significantly across databases, with higher rates than ECDC TESSy suggesting false positives when including unspecified pertussis concepts.
- In SIDIAP, including discharge diagnoses increased the estimated IR to 45.0 per 100,000 person-years.
- The estimated sensitivity and PPV for BorPer cases using combined PC diagnoses in SIDIAP were approximately 85% and 72%, respectively.
Conclusions:
- Using only specific BorPer concepts in PC databases leads to low sensitivity, while including unspecified concepts introduces false positives (estimated ~28%).
- An estimated 15% of cases may not be captured in PC databases if they are exclusively diagnosed in hospitals.
- Quantifying algorithm impact across databases and benchmarking with surveillance data provides valuable approximate estimates of algorithm validity for vaccine safety.
More Related Videos
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding in Epidemiological Studies
Bias in Epidemiological Studies
Statistical Methods for Analyzing Epidemiological Data
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...

