Related Experiment Video
Updated: Jan 12, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Training a Smoking Status Probabilistic Model Using Cotinine Levels in a Large Claims Database
Dominique Medaglio1,2, Charles E Leonard1,2,3, Alisa J Stephens Shields1,2,3,4
1Center for Pharmacoepidemiology Research and Training, Center for Clinical Epidemiology and Biostatistics, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA.
Introduction:
Smoking status is an important confounder for many epidemiologic studies, yet it is not well documented in common sources of real-world data, including administrative claims. Probabilistic models can be used to create a proxy for smoking status, yet most published models have been trained using self-reported data. The objective of this study was to train a smoking status probabilistic model using cotinine values available in a large claims database.
Methods:
Beneficiaries were included if they had at least one cotinine measurement and were categorized as a "current smoker" if their serum or plasma cotinine value was ≥5 ng/mL or urine cotinine value was ≥30 ng/mL. Predictors were collected across one year prior to the cotinine assessment date. The model was fit using logistic regression with stepwise forward selection. Model performance was assessed using discrimination and calibration.
Results:
The final model yielded an area under the receiver operating characteristic curve of 0.77 (95%CI:0.75-0.78) and was well calibrated across most prediction deciles. The strongest predictors included diagnosis codes for smoking and drug abuse, and number of medications. The model was found to be highly specific, yet not sensitive at probability cutoffs ≥0.2.
Conclusions:
A smoking status model was developed and internally validated for application in claims data, using available cotinine values to define smoking status and found to have acceptable discrimination and calibration. The model is based on 26 predictors, fewer than other similar published smoking status models. External validation of the model should be a next step toward utilizing the model for epidemiological research.
Implications:
This study tests the utility of cotinine values to validate a smoking status probabilistic model, which has not been done in the literature to date. The results were robust to various cotinine levels used to define smoking status, per current guidance. The final model uses only 26 factors to predict smoking status, simplifying the application of the model in other claims databases.
More Related Videos
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Observational Studies
There are three types of observational studies – Prospective, retrospective, and cross-sectional.
Prospective Study
Prospective studies, also known as longitudinal or cohort studies, are carried out by collecting future data from groups sharing similar characteristics. One...