Using natural language processing to identify opioid use disorder in electronic health record data
Jade Singleton1, Chengxi Li2, Peter D Akpunonu3
1Department of Epidemiology, College of Public Health, University of Kentucky, Lexington, KY 40536, United States; University of Kentucky Healthcare IT Department, Business Intelligence, Lexington, KY 40517, United States.
Natural Language Processing (NLP) combined with ICD-10-CM codes improves opioid use disorder (OUD) detection in electronic health records (EHR). This approach increases OUD prevalence estimates by 29.5% compared to using codes alone, enhancing epidemiological surveillance.
Area of Science:
- Public Health
- Health Informatics
- Epidemiology
Background:
- Rising opioid prescriptions correlate with increased opioid use disorder (OUD) and adverse outcomes.
- Accurate epidemiological surveillance of OUD is crucial for effective prevention strategies but faces challenges.
- Electronic health records (EHR) offer a potential data source for OUD surveillance.
Purpose of the Study:
- To compare the effectiveness of two methods for identifying OUD in EHR data: natural language processing (NLP) and ICD-10-CM diagnostic codes.
- To ascertain the prevalence of OUD using both NLP and ICD-10-CM codes.
- To evaluate the contribution of NLP in identifying OUD cases missed by diagnostic codes.
Main Methods:
- EHR data from hospital and emergency department visits between 2017-2019 were analyzed.
- A rule-based NLP algorithm was developed and refined using a stepwise process with EHR data.
- ICD-10-CM discharge codes were extracted, and NLP was applied to unstructured clinical notes.
Main Results:
- A combined approach using NLP and ICD-10-CM codes identified 2,332 unique OUD cases, a 29.5% increase over ICD-10-CM codes alone (6.1% vs. 7.9% prevalence).
- NLP identified 521 OUD cases (22.3%) not captured by ICD-10-CM codes, while ICD-10-CM codes identified 430 cases (18.4%) missed by NLP.
- The NLP algorithm demonstrated high accuracy with an estimated sensitivity of 81.8% and specificity of 97.5% compared to expert manual review.
Conclusions:
- NLP-based algorithms effectively automate data extraction and identify OUD from unstructured EHR data.
- Combining NLP with ICD-10-CM codes provides the most comprehensive ascertainment of OUD in EHR.
- NLP should be integrated into epidemiological studies utilizing EHR data for improved OUD surveillance.
More Related Videos
Related Concept Videos
Opioid Analgesics: Morphine and Other Natural Cogeners
Opioid Receptors: Overview
Opioid Analgesics: Synthetic and Semisynthetic Opioids
Analgesia and Pain Management
Prescription, Nonprescription and Orphan Drugs
The misuse and addiction to prescription drugs is a growing problem that can affect people of all age groups, specifically teenagers. This can happen when prescription medications are used in ways not intended by the prescriber, such as taking someone else's prescription or using medication for...
Drug Abuse and Addiction: Pharmacological Phenomena


