Related Experiment Video
Updated: Aug 6, 2026

High-throughput and Comprehensive Drug Surveillance Using Multisegment Injection-Capillary Electrophoresis-Mass Spectrometry
Published on: April 23, 2019
Natural Language Processing to Identify Substance Use in Electronic Health Records: A Scoping Review
Chidimma Doris Azubuike1, Ahmed Farrag1, Kimia Zandbiglari1
1Department of Pharmaceutical Outcomes and Policy, University of Florida, Gainesville, FL, USA.
Background:
This scoping review aimed to characterize natural language processing (NLP) techniques deployed for identifying substance use in electronic health records (EHRs) and to compare the performance of these techniques by substance type.
Methods:
We conducted a systematic search of PubMed, the Cochrane Library, Embase, Web of Science, ACM Digital Library, IEEE Xplore, and Scopus for peer-reviewed original research published in English. Studies were eligible if they applied NLP to identify non-prescription or problematic substance use in EHRs, provided full-text access, and reported quantitative performance metrics.
Results:
A total of 86 studies met the inclusion criteria. NLP-identified substance use included tobacco (n = 42), alcohol (n = 26), non-prescription opioids (n = 30), cannabinoids (n = 9), stimulants (n = 6), and polysubstance use and/or other drugs (n = 16). NLP techniques included rule-based (n = 46), conventional machine learning (n = 39), deep learning (n = 17), and large language/transformer-based models (n = 22), with some studies applying multiple techniques. Annotation guidelines were available for 30 studies, and only 22 published their codes. Most studies reported performance metrics exceeding 0.80.
Conclusions:
Tobacco, opioids, and alcohol were the most frequently identified substances, whereas stimulants and cannabinoids were markedly underrepresented. Across the reviewed literature, transparency and reproducibility were limited, with few studies publishing code or detailed model specifications. Nevertheless, reported performance metrics were generally high. NLP techniques showed high performance in identifying tobacco, opioids, and alcohol use in EHRs, but stimulants and cannabinoids remain underrepresented. Transparency and reproducibility remain limited, underscoring the need for routine sharing of code, datasets, and model specifications.
Related Concept Videos
Substance Use Disorders Affecting Sleep
Understanding the concepts of physical dependence,...
Pharmaceutical Poisoning: Potential Scenarios
Pharmacovigilance
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
Drug Discovery: Overview