Related Experiment Video
Updated: Aug 6, 2026

High-throughput and Comprehensive Drug Surveillance Using Multisegment Injection-Capillary Electrophoresis-Mass Spectrometry
Published on: April 23, 2019
Natural Language Processing to Identify Substance Use in Electronic Health Records: A Scoping Review
Chidimma Doris Azubuike1, Ahmed Farrag1, Kimia Zandbiglari1
1Department of Pharmaceutical Outcomes and Policy, University of Florida, Gainesville, FL, USA.
Natural Language Processing (NLP) effectively identifies tobacco, opioid, and alcohol use in electronic health records (EHRs). However, stimulant and cannabinoid use detection needs improvement, and transparency in NLP methods is limited.
Area of Science:
- Computational linguistics
- Health informatics
- Substance use research
Background:
- Electronic health records (EHRs) contain valuable data on substance use.
- Natural Language Processing (NLP) offers potential for identifying substance use within unstructured EHR data.
- Characterizing NLP techniques and their performance in substance use identification is crucial for advancing clinical informatics.
Purpose of the Study:
- To systematically review and characterize Natural Language Processing (NLP) techniques used for identifying substance use in electronic health records (EHRs).
- To compare the performance of different NLP techniques across various substance types.
- To assess the transparency and reproducibility of NLP methods applied to substance use identification in EHRs.
Main Methods:
- Conducted a systematic literature search across major scientific databases (PubMed, Cochrane Library, Embase, Web of Science, ACM Digital Library, IEEE Xplore, Scopus).
- Included peer-reviewed original research in English applying NLP to identify non-prescription or problematic substance use in EHRs with full-text access and quantitative performance metrics.
- Extracted data on NLP techniques, substance types, and performance metrics from 86 eligible studies.
Main Results:
- Tobacco (42 studies), non-prescription opioids (30 studies), and alcohol (26 studies) were the most frequently identified substances.
- NLP techniques included rule-based (46 studies), conventional machine learning (39 studies), deep learning (17 studies), and large language/transformer-based models (22 studies).
- Most studies reported high performance metrics (>0.80), but transparency was limited, with only 30 studies providing annotation guidelines and 22 sharing code.
Conclusions:
- NLP demonstrates high performance in identifying common substance uses like tobacco, opioids, and alcohol in EHRs.
- Stimulants and cannabinoids remain underrepresented in NLP-based substance use identification research.
- Limited transparency and reproducibility necessitate routine sharing of code, datasets, and model specifications for advancing NLP in EHR research.
Related Concept Videos
Substance Use Disorders Affecting Sleep
Understanding the concepts of physical dependence,...
Pharmaceutical Poisoning: Potential Scenarios
Pharmacovigilance
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
Drug Discovery: Overview