Related Experiment Video
Updated: Nov 13, 2025

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Using computable knowledge mined from the literature to elucidate confounders for EHR-based pharmacovigilance
Scott A Malec1, Peng Wei2, Elmer V Bernstam3
1University of Pittsburgh School of Medicine, Department of Biomedical Informatics, Pittsburgh, PA, United States.
This study shows that using literature-derived confounders improves causal inference in drug safety research. Semantic vector search was better than string search for reducing confounding bias in observational data.
Area of Science:
- Pharmacovigilance
- Biomedical Informatics
- Causal Inference
Background:
- Drug safety studies often use observational data, which is susceptible to confounding bias.
- Identifying confounders requires deep knowledge of complex biological systems, posing a challenge for reliable causal inference.
- Literature-derived computable knowledge may offer a solution for identifying potential confounders.
Purpose of the Study:
- To test if incorporating literature-derived confounders improves causal inference from observational data in drug safety research.
- To introduce and evaluate methods for searching biomedical literature for confounder candidates.
Main Methods:
- Developed and compared two methods: semantic vector-based and string-based confounder search using SemMedDB.
- Queried SemMedDB for confounders by linking drug indications (exposure) to adverse events (outcomes).
- Integrated literature-derived confounders into statistical and causal models using NLP-processed clinical notes and evaluated against known drug-adverse event relationships.
Main Results:
- Semantic vector-based search demonstrated superiority in reducing confounding bias compared to string-based search.
- The impact of the quantity of literature-derived confounders on bias reduction was inconclusive.
- Naive association measures (chi-squared, reporting odds ratio) were used for comparison.
Conclusions:
- Semantic search methods are promising for enhancing confounder identification in drug safety.
- Further research is needed to optimize the use of literature-derived confounders, especially regarding quantity.
- Consideration of targeted learning estimation and expert adjudication is recommended for complex covariate scenarios.
Related Concept Videos
Pharmacovigilance
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
Analysis of Population Pharmacokinetic Data
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
Pharmacokinetic Models: Overview
There are three primary types of models: empirical, compartment, and physiological. Empirical models, with minimal...

