Related Experiment Video
Updated: Mar 27, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Identification of smoking using Medicare data--a validation study of claims-based algorithms
Rishi J Desai1, Daniel H Solomon1,2, Nancy Shadick2
1Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital & Harvard Medical School, Boston, MA, USA.
Claims-based algorithms can identify smokers, but with limited sensitivity. However, they offer very high specificity, making them cautiously useful for observational studies when other methods are unavailable.
Area of Science:
- Health Informatics
- Epidemiology
- Medical Record Analysis
Background:
- Accurate identification of smoking status is crucial for epidemiological research and clinical studies.
- Self-reported smoking data is the gold standard but can be subject to recall bias.
- Claims data offers a potentially scalable alternative for identifying smoking behavior in large patient populations.
Purpose of the Study:
- To evaluate the accuracy of Medicare claims-based algorithms in identifying current smokers.
- To compare algorithms using diagnosis/procedure codes versus those including anti-smoking prescriptions.
- To assess algorithm performance using different look-back periods for claims data.
Main Methods:
- Two claims-based algorithms were developed to identify smoking status in Medicare beneficiaries.
- Algorithm 1 utilized only diagnosis and procedure codes; Algorithm 2 included these plus anti-smoking prescriptions.
- Performance was assessed against self-reported smoking data (gold standard) using sensitivity, specificity, NPV, PPV, and AUC.
Main Results:
- Claims-based algorithms demonstrated high specificity (100%) but low sensitivity (ranging from 9.8% to 27.9%).
- Incorporating pharmacy claims and using all available pre-index data improved sensitivity, NPV, and AUC compared to limited data.
- The algorithm using diagnosis/procedure codes and all pre-index claims had an AUC of 0.64.
Conclusions:
- Claims-based algorithms can identify smokers with high specificity, though sensitivity remains a limitation.
- These algorithms may be cautiously considered for identifying smoking in observational studies when self-report data is unavailable.
- Further refinement of claims-based algorithms could enhance their utility in epidemiological research.
More Related Videos
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Chronic Obstructive Pulmonary Disease-IV: Assessement and Diagnostic Studies
Medical History
Observational Studies
There are three types of observational studies – Prospective, retrospective, and cross-sectional.
Prospective Study
Prospective studies, also known as longitudinal or cohort studies, are carried out by collecting future data from groups sharing similar characteristics. One...

