Related Experiment Video
Updated: Mar 9, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Development of an algorithm for determining smoking status and behaviour over the life course from UK electronic
Mark D Atkinson1, Jonathan I Kennedy2, Ann John2
1Farr Institute, Swansea University Medical School, Swansea, SA2 8PP, UK. M.Atkinson@swansea.ac.uk.
Background:
Patients' smoking status is routinely collected by General Practitioners (GP) in UK primary health care. There is an abundance of Read codes pertaining to smoking, including those relating to smoking cessation therapy, prescription, and administration codes, in addition to the more regularly employed smoking status codes. Large databases of primary care data are increasingly used for epidemiological analysis; smoking status is an important covariate in many such analyses. However, the variable definition is rarely documented in the literature.
Methods:
The Secure Anonymised Information Linkage (SAIL) databank is a repository for a national collection of person-based anonymised health and socio-economic administrative data in Wales, UK. An exploration of GP smoking status data from the SAIL databank was carried out to explore the range of codes available and how they could be used in the identification of different categories of smokers, ex-smokers and never smokers. An algorithm was developed which addresses inconsistencies and changes in smoking status recording across the life course and compared with recorded smoking status as recorded in the Welsh Health Survey (WHS), 2013 and 2014 at individual level. However, the WHS could not be regarded as a "gold standard" for validation.
Results:
There were 6836 individuals in the linked dataset. Missing data were more common in GP records (6%) than in WHS (1.1%). Our algorithm assigns ex-smoker status to 34% of never-smokers, and detects 30% more smokers than are declared in the WHS data. When distinguishing between current smokers and non-smokers, the similarity between the WHS and GP data using the nearest date of comparison was κ = 0.78. When temporal conflicts had been accounted for, the similarity was κ = 0.64, showing the importance of addressing conflicts.
Conclusions:
We present an algorithm for the identification of a patient's smoking status using GP self-reported data. We have included sufficient details to allow others to replicate this work, thus increasing the standards of documentation within this research area and assessment of smoking status in routine data.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
11:21Methodology for Establishing a Community-Wide Life Laboratory for Capturing Unobtrusive and Continuous Remote Activity and Health Data
Published on: July 27, 2018
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Observational Studies
There are three types of observational studies – Prospective, retrospective, and cross-sectional.
Prospective Study
Prospective studies, also known as longitudinal or cohort studies, are carried out by collecting future data from groups sharing similar characteristics. One...
Chronic Obstructive Pulmonary Disease-IV: Assessement and Diagnostic Studies
Medical History
Lifestyle Factors and Health
Benefits of Physical Activity
Physical activity, whether through structured exercise or casual activities like walking, biking, or dancing, is a cornerstone of a...
Physical Assessment of the Respiratory Tract I: Health History
Subjective Data
Subjective data provides vital information about the patient's health history and symptoms. This data is typically collected through interviews in which patients describe their experiences, symptoms, and concerns.
Health history and...
Longitudinal Research