Related Experiment Video
Updated: Apr 15, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Validation of Diagnostic Groups Based on Health Care Utilization Data Should Adjust for Sampling Strategy
Geneviève Cadieux1, Robyn Tamblyn, David L Buckeridge
1*Dalla Lana School of Public Health, University of Toronto, Toronto, ON †Department of Epidemiology, Biostatistics and Occupational Health, McGill University ‡Direction de la Santé Publique de Montréal §Department of Medicine, McGill University, Montreal, QC, Canada.
Accurate measurement of disease prevalence using health care data requires stratified sampling. This study proposes a method to improve the positive predictive value (PPV) and accuracy of diagnostic group measurements, especially for low-prevalence conditions.
Area of Science:
- Health Informatics
- Biostatistics
- Population Health Surveillance
Background:
- Validating health outcomes from healthcare utilization data is crucial for learning health systems.
- Existing validation methods often lack awareness of stratified sampling needs for complex diagnostic groups.
- Low-prevalence diagnostic codes present unique challenges in accurate measurement.
Purpose of the Study:
- To propose a novel statistical method for validating diagnostic group measurements.
- To address challenges posed by varying prevalences and low prevalence of diagnostic codes.
- To enhance the accuracy of outcome measurement in healthcare data.
Main Methods:
- Oversampling low-prevalence diagnostic codes to estimate the positive predictive value (PPV) as a weighted average.
- Oversampling claims within low-prevalence diagnostic groups for bias-adjusted sensitivity and specificity estimation.
- Utilizing a stratified sampling design and corresponding statistical methods.
Main Results:
- The proposed method improves the accuracy of positive predictive value (PPV) estimation for diagnostic groups.
- Bias-adjusted estimators for sensitivity and specificity are generated, accounting for oversampling.
- Illustrative example demonstrates application in acute respiratory illness surveillance.
Conclusions:
- Failure to account for diagnostic code prevalence underestimates PPV due to false positives in low-prevalence codes.
- Improper adjustment for oversampling leads to overestimated sensitivity and underestimated specificity.
- The proposed method offers a robust approach to validating complex diagnostic groups in healthcare data.
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...

