Related Experiment Video
Updated: Dec 5, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Extracting medication information from unstructured public health data: a demonstration on data from population-based
Robert Chen1,2, Joyce C Ho1,3, Jin-Mann S Lin4
1Chronic Viral Diseases Branch, Division of High-Consequence Pathogens and Pathology, National Center for Emerging and Zoonotic Infectious Diseases, Centers for Disease Control and Prevention, 1600 Clifton Rd. NE, Mailstop H24-12, Atlanta, GA, 30329, USA.
This study introduces an automated framework using natural language processing (NLP) to extract and process unstructured clinical data, significantly reducing manual labor for medication and reason mapping in epidemiological studies.
Area of Science:
- Clinical Epidemiology
- Natural Language Processing
- Data Science
Background:
- Unstructured data from clinical studies is valuable but difficult to analyze manually.
- Manual data extraction is time-consuming, labor-intensive, and prone to errors.
- Automation is needed to efficiently process large volumes of unstructured clinical data.
Purpose of the Study:
- To develop and demonstrate an automation framework for extracting and processing unstructured clinical data.
- To reduce the manual effort required for data analysis in epidemiological studies.
- To improve the efficiency and accuracy of data processing for medication and reasons for use.
Main Methods:
- Utilized two natural language processing (NLP) tools for medication and reason extraction.
- Applied spell-checking and mapped medication names to generic and Anatomical Therapeutic Chemical (ATC) classifications.
- Processed reasons for medication using the Lancaster stemmer and mapped them to disease classes based on organ systems.
- Demonstrated the framework on Myalgic Encephalomyelitis/Chronic Fatigue Syndrome (ME/CFS) data.
Main Results:
- Successfully condensed 1266 distinct medication names into 89 ATC categories.
- Reduced 1432 distinct reasons for medication use into 65 categories using NLP.
- Automation reduced manual mapping effort by 84.4% for medications and 59.4% for reasons.
- Improved the precision of mapped results compared to manual processing.
Conclusions:
- The NLP-based automation framework effectively processes unstructured clinical data.
- The method is adaptable for less established databases and can incorporate new knowledge sources.
- Condensing features into interpretable categories enhances subsequent machine learning and data mining studies.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
10:17High-throughput and Comprehensive Drug Surveillance Using Multisegment Injection-Capillary Electrophoresis-Mass Spectrometry
Published on: April 23, 2019
Related Concept Videos
Analysis of Population Pharmacokinetic Data
Dosage Regimens: Partial Pharmacokinetic Parameters
Statistical Methods for Analyzing Epidemiological Data
Bioavailability Study Design: Healthy Subjects Versus Patients
Data Collection I
Mechanistic Models: Compartment Models in Individual and Population Analysis