Related Experiment Video
Updated: May 27, 2025

Impact Assessment of Repeated Exposure of Organotypic 3D Bronchial and Nasal Tissue Culture Models to Whole Cigarette Smoke
Published on: February 12, 2015
Comparative ranking of marginal confounding impact of natural language processing-derived versus structured features
Joseph M Plasek1, Richard D Wyss2, Janick G Weberpals2
1Division of General Internal Medicine, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA.
Objective:
To explore the ability of natural language processing (NLP) methods to identify confounder information beyond what can be identified using claims codes alone for pharmacoepidemiology.
Methods:
We developed a retrospective cohort for high vs low dose proton pump inhibitors from linked Medicare claims (2008-2017) and electronic health record data for patients with a history of peptic ulcer disease or osteoarthritis. Clinical notes authored one year before first dispensing date were processed by off-the-shelf tools: bag-of-n-grams, latent Dirichlet allocation, a linguistics-focused tool, BERT sentence embeddings, BioBERT word embeddings, and GloVe word embeddings. Candidate features were ranked using Bross formula, a simple way to rank the marginal confounding impact of binary features on estimated causal effects.
Results:
The marginal confounding impact in the Bross rankings of NLP-derived features trended from 39 % in the top 100 to 77 % in the top 500 to 93 % in the top 5000 among patients with peptic ulcer disease. More specifically, the top 25 confounders are largely from factors identified by domain experts and structured fields, and the marginal impact of these confounders is stronger than others. Features 25 to 50 include features identified by a linguistics-focused tool and embeddings, whereas features 50 to 100 include more embeddings and bag-of-ngrams. After 100, the curve flattens, meaning that the marginal impact of those potential confounders gets smaller. Similarly, among patients with osteoarthritis, NLP-derived features trended from 66 % in the top 100 to 84 % in the top 500 to 95 % in the top 5000 when the outcome was gastrointestinal bleed and from 47 % in the top 100 to 81 % in the top 500 to 94 % in the top 5000 when the outcome was acute kidney injury. Similar trends were observed in the information gain data, though NLP-derived features had higher baselines.
Conclusions:
NLP contributed to finding large numbers of features that can supplement claims data and prespecified variables to help provide additional confounder information. We found that unsupervised off-the-shelf NLP tools can scale to generate large numbers of features appropriate for high-dimensional proxy adjustment and pharmacoepidemiology use cases.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
09:35A Protocol for Using Gene Set Enrichment Analysis to Identify the Appropriate Animal Model for Translational Research
Published on: August 16, 2017
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Analysis of Population Pharmacokinetic Data
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding in Epidemiological Studies
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
Mechanistic Models: Compartment Models in Individual and Population Analysis