Related Experiment Video
Updated: Aug 9, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Unstructured Data Are Superior to Structured Data for Eliciting Quantitative Smoking History From the Electronic
John C Ruckdeschel1,2, Mark Riley1, Sriram Parsatharathy1
1MetiStream, Inc, Vienna, VA.
This study developed a method using natural language processing (NLP) to extract smoking history from clinical notes. This approach accurately identifies patients eligible for lung cancer screening with low-dose computed tomography (LDCT).
Area of Science:
- Medical informatics
- Computational linguistics
- Public health
Background:
- Accurate identification of smoking status is crucial for lung cancer screening eligibility.
- Traditional structured data often lacks detailed smoking history, hindering cohort identification for low-dose computed tomography (LDCT).
- Natural language processing (NLP) offers a potential solution for extracting nuanced clinical information from unstructured text.
Purpose of the Study:
- To develop and validate a method for extracting smoking status and quantitative smoking history from clinician notes.
- To facilitate the identification of patient cohorts eligible for LDCT screening for early lung cancer detection.
- To improve the accuracy and efficiency of determining LDCT eligibility based on smoking criteria.
Main Methods:
- Utilized natural language processing (NLP) and named entity recognition on clinician notes from the MIMIC-III database.
- Developed algorithms to extract quantitative smoking history, including pack-years and quit dates.
- Validated extracted data against manual chart reviews, calculating an F-score for accuracy.
Main Results:
- NLP identified 1,930 ever-smokers from clinician notes, a significant increase compared to structured data.
- The method successfully identified 276 patients eligible for LDCT screening based on smoking and age criteria, using USPSTF guidelines.
- The F-score for identifying eligible patients using this NLP-based approach was 0.88, demonstrating high accuracy.
Conclusions:
- Unstructured clinical data, when processed by NLP, can accurately identify specific patient cohorts for LDCT screening.
- This NLP-driven method enhances the precision of cohort identification for lung cancer early detection programs.
- The findings support the integration of NLP tools in electronic health records to optimize screening eligibility determination.
More Related Videos
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Methods of Documentation VII: EMR
Data Collection II
Observational Studies
There are three types of observational studies – Prospective, retrospective, and cross-sectional.
Prospective Study
Prospective studies, also known as longitudinal or cohort studies, are carried out by collecting future data from groups sharing similar characteristics. One...
Data Reporting and Recording
Physical Assessment of the Respiratory Tract I: Health History
Subjective Data
Subjective data provides vital information about the patient's health history and symptoms. This data is typically collected through interviews in which patients describe their experiences, symptoms, and concerns.
Health history and...

